# Omni Language AI Research (OLAResearch) Omni Language AI Research (OLAResearch) develops scalable and multimodal AI systems tailored for real-world applications. The research collective operates at the intersection of deep learning, linguistics, and healthcare, with a strong focus on bridging the gap between high-resource and low-resource languages, ensuring equitable access to advanced AI. ## Core Information - **Domain**: https://www.olaresearch.org - **GitHub**: https://github.com/OLAResearch - **Hugging Face**: https://huggingface.co/OLAResearchX - **Google Groups**: https://groups.google.com/g/omni-language-ai-research - **Affiliations**: TurkuNLP, ELLIS Institute Finland, University of Turku --- ## Research Directions ### [FINe-Health Foundry](https://www.olaresearch.org/fine-health-foundry.html) Developing a nationwide, versatile foundation health model using Finland's national health databases for clinical decision support, individual disease prevention, and policy simulations. In collaboration with ELLIS Institute Finland, Aalto University, University of Helsinki, and University of Turku. ### [AI for Health](https://www.olaresearch.org/ai4health.html) Developing next-generation AI systems for pathology, treatment optimization, and automated clinical decision-making. ### [AI for Mental Health](https://www.olaresearch.org/ai-mh.html) Advancing automated mental health analysis through domain-specific LLMs, emotional analysis, and resource-driven modeling. ### [Multilingual NLP](https://www.olaresearch.org/multilingual-nlp.html) Multilingual NLP for low-resource languages, ensuring global applicability of language technology. ### [NLP for Health (Archived)](https://www.olaresearch.org/nlp4health.html) Advancing clinical information management through automated medical coding, patient outcome prediction, and biomedical reasoning. --- ## Selected Publications ### 2026 Publications * **[MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models](https://www.olaresearch.org/MaLA/)** * *Authors*: Shaoxiong Ji, Zihao Li, Jaakko Paavola, Peiqin Lin, Pinzhen Chen, Dayyán O'Brien, Hengyu Luo, Hinrich Schütze, Jörg Tiedemann, Barry Haddow * *Venue*: COLM 2026 * *Links*: [Paper URL](https://www.olaresearch.org/MaLA/mala-corpus.pdf) | [Code](https://github.com/MaLA-LM/emma-500) | [Data](https://huggingface.co/collections/MaLA-LM/mala-corpus-66e05127641a51de34d39529) | [Model](https://huggingface.co/MaLA-LM/emma-500-llama2-7b) * *Summary*: Introduces the MaLA suite: a 74B-token corpus across 939 languages, a balanced 136B training mix, and the EMMA-500 model optimized for cross-lingual transfer. * **[Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?](https://www.olaresearch.org/LingReason/)** * *Authors*: Renhao Pei, Yihong Liu, Sampo Pyysalo, Hinrich Schuetze, Shaoxiong Ji * *Venue*: Preprint / arXiv 2606.03782 * *Links*: [Paper URL](https://arxiv.org/abs/2606.03782) | [Code](https://github.com/OLAResearch/LingReason) | [Dataset](https://huggingface.co/datasets/OLAResearchX/LingReason) * *Summary*: Proposes a pipeline for automatically generating step-by-step linguistic reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar-rule banks to improve low-resource machine translation. * **[Cross-Model Memory Transfer via Target-Side Reader Adaptation](https://www.olaresearch.org/XMemTransfer/)** * *Authors*: Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji * *Venue*: Preprint * *Links*: [Paper URL](https://www.olaresearch.org/XMemTransfer/paper.pdf) | [Code](https://github.com/OLAResearch/XMemTransfer) | [Models](https://huggingface.co/collections/OLAResearchX/xmemtransfer) * *Summary*: Investigates the portability of external memory structures (Engram) across different model backbones and tokenizers via target-side reader adaptation. * **[Test-Time Scaling of Reasoning Models for Machine Translation](https://www.olaresearch.org/TTS4MT/)** * *Authors*: Zihao Li, Shaoxiong Ji, Jörg Tiedemann * *Venue*: EACL 2026 (Long Papers) * *Links*: [ACL Anthology](https://aclanthology.org/2026.eacl-long.133/) * *Summary*: Investigates the efficacy of test-time scaling (inference-time computation) in machine translation across diverse benchmarks. * **[Data-Centric Continual Pre-training for 500+ Languages: A New Bilingual Translation Corpus and Multilingual Models](https://www.olaresearch.org/EMMA-500-Gen2/)** * *Authors*: Shaoxiong Ji, Zihao Li, Jaakko Paavola, Hengyu Luo, Jörg Tiedemann * *Venue*: ACL Findings 2026 * *Links*: [ACL Anthology](https://aclanthology.org/2026.findings-acl.937/) | [Code](https://github.com/MaLA-LM/emma-500) | [Data](https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus) | [Models](https://huggingface.co/collections/MaLA-LM/emma-500) * *Summary*: Investigates the impact of bilingual translation data for massively multilingual continual pre-training of the Llama 3 family of models to 500 languages, releasing the MaLA bilingual corpus and the EMMA Llama 3 models trained up to 671B tokens. * **You Never Know a Person, You Only Know Their Defenses: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations** * *Authors*: Hongbin Na, Zimu Wang, Zhaoming Chen, Peilin Zhou, Yining Hua, Grace Ziqi Zhou, Haiyang Zhang, Tao Shen, Wei Wang, John Torous, Shaoxiong Ji, Ling Chen * *Venue*: ACL 2026 Findings * *Links*: [ACL Anthology](https://aclanthology.org/2026.findings-acl.708/) * **Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations** * *Authors*: Hongbin Na, Zimu Wang, Zhaoming Chen, Yining Hua, Rena Gao, Kailai Yang, Ling Chen, Wei Wang, Shaoxiong Ji, John Torous, Sophia Ananiadou * *Venue*: BioNLP 2026 * *Links*: [ACL Anthology](https://aclanthology.org/2026.bionlp-1.75/) * **A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding** * *Authors*: Dilara Torunoğlu-Selamet et al. (including Shaoxiong Ji) * *Venue*: LREC 2026 * *Links*: [DOI Link](https://doi.org/10.63317/5cvnbcoktfo2) * **Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning** * *Authors*: Shuo Yu, Yingbo Wang, Ruolin Li, Guchun Liu, Yanming Shen, Shaoxiong Ji, Bowen Li, Fengling Han, Xiuzhen Zhang, Feng Xia * *Venue*: ACM Transactions on Information Systems (TOIS 2026) * *Links*: [DOI Link](https://doi.org/10.1145/3786600) * **XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad** * *Authors*: Mohsinul Kabir, Tasnim Ahmed, Md Mezbaur Rahman, Shaoxiong Ji, Hassan Alhuzali, Yuechen Jiang, Jimin Huang, Sophia Ananiadou * *Venue*: EMNLP 2026 * *Links*: [Paper URL](https://arxiv.org/abs/2601.14063) * **Psychologically-Grounded Graph Modeling for Interpretable Depression Detection** * *Authors*: Rishitej Reddy Vyalla, Kritarth Prasad, Avinash Anand, Erik Cambria, Shaoxiong Ji, Faten S Alamri, Zhengkui Wang * *Venue*: arXiv 2604.24126 * **Model-Based Quality Assessment for Massively Multilingual Parallel Data** * *Authors*: Abdelaziz Ibrahim, Zihao Li, Jörg Tiedemann, Shaoxiong Ji * *Venue*: arXiv 2606.00285 * **Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance** * *Authors*: Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jiancheng Yang * *Venue*: EMNLP Findings 2026 * *Links*: [Paper URL](https://arxiv.org/abs/2606.18613) ### 2025 Publications * **Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources** * *Authors*: Zihao Li, Shaoxiong Ji, Hengyu Luo, Jörg Tiedemann * *Venue*: Conference on Language Modeling (COLM 2025) * *Links*: [OpenReview](https://openreview.net/pdf?id=mpTIzK4Zca) * **GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models** * *Authors*: Hengyu Luo, Zihao Li, Joseph Attieh et al. (including Shaoxiong Ji, Jörg Tiedemann) * *Venue*: EMNLP 2025 System Demonstrations * *Links*: [ACL Anthology](https://aclanthology.org/2025.emnlp-demos.43/) * **Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires** * *Authors*: Doan Nam Long Vu, Rui Tan, Lena Moench, Svenja Jule Francke, Daniel Woiwod, Florian Thomas-Odenthal, Sanna Stroth, Tilo Kircher, Christiane Hermann, Udo Dannlowski, Hamidreza Jamalabadi, Shaoxiong Ji * *Venue*: arXiv 2510.25384 * *Links*: [Paper URL](https://arxiv.org/abs/2510.25384) --- ## Core Team Members * **[Shaoxiong Ji](https://www.olaresearch.org/jis/)** (Principal Investigator) * *Affiliations*: ELLIS Institute Finland & University of Turku * *Interests*: NLP & AI for Health * **[Zihao Li](https://www.zihao.cool)** (Doctoral Researcher) * *Affiliations*: University of Helsinki (Co-supervision with Jörg Tiedemann) * *Interests*: Multilingual NLP * **[Renhao Pei](https://peirh.github.io)** (Doctoral Researcher) * *Affiliations*: ELLIS Institute Finland & University of Turku * *Interests*: Multilingual NLP * **[Mingyuan Li](https://www.utu.fi/en/people/mingyuan-li)** (Postdoctoral Researcher) * *Affiliations*: ELLIS Institute Finland & University of Turku * *Interests*: Machine Learning * **[Md Mohsinul Kabir](https://scholar.google.com/citations?user=eVVCkREAAAAJ&hl=en)** (ELLIS PhD) * *Affiliations*: University of Manchester (co-advised with Sophia Ananiadou) * *Interests*: NLP * **[Yicong Wu](https://www.utu.fi/fi/ihmiset/yicong-wu)** (Doctoral Researcher) * *Affiliations*: ELLIS Institute Finland & University of Turku * *Interests*: AI for Health * **[Naveen Vakada](https://scholar.google.com/citations?user=8oHYRMcAAAAJ&hl=en)** (Doctoral Researcher) * *Affiliations*: ELLIS Institute Finland & University of Turku * *Interests*: Multimodal AI * **[Zhibo Man](https://www.utu.fi/en/people/zhibo-man)** (Postdoctoral Researcher) * *Affiliations*: ELLIS Institute Finland & University of Turku * *Interests*: NLP