One year ago, I moved my research group to ELLIS Institute Finland and the University of Turku, after a year as a research group leader at the Technical University of Darmstadt. Today marks exactly one year since the move. Over the past year, as an ELLIS PI, I have broadened the scope of our research. We focused on three areas: multilingual LLMs, LLM reasoning and memory, and AI for health and mental health. We published papers at ACL, COLM, EMNLP and NeurIPS, started a national foundation health model project, released a trillion-token parallel corpus, and ran a shared task and a workshop. This post looks back on what the group did with our collaborators, and on the people who made it happen.
Research
Counting papers is a bad practice, but it is hard to avoid in an end-of-year review like this one. Across 2025 and 2026 I co-authored 22 papers. Fifteen of them are accepted or published at peer-reviewed venues, and the other seven are preprints currently under review. The full list is on the publications page. The work falls into three connected lines.
Multilingual large language models
This is the longest-running line, and the one I wrote about in The MaLA-LM Journey. After several rounds of review, two papers finally landed this year. The first is MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models at COLM 2026. The second is Data-Centric Continual Pre-training for 500+ Languages (project page) in ACL 2026 Findings. They join Rethinking Multilingual Continual Pretraining (COLM 2025), in which we studied data-mixing strategies for multilingual continual pretraining, and the GlotEval evaluation suite (EMNLP 2025 System Demonstrations), which we built to enable massively multilingual evaluation.
Translation for low-resource languages was a recurring theme. Two papers asked whether reasoning helps translation: Test-Time Scaling of Reasoning Models for Machine Translation (project page) (EACL 2026) and Reasoning over Grammar (project page), which tests whether synthetic linguistic reasoning traces help low-resource MT (EMNLP 2026 Findings). On the cultural side, we contributed to XCR-Bench (EMNLP 2026) and CulturALL (AACL-IJCNLP 2026), two benchmarks for cross-cultural competence in LLMs, and to a cross-lingual benchmark for multimodal idiomaticity at LREC 2026.
LLM reasoning and memory
This line grew quickly this year. Cross-Model Memory Transfer via Target-Side Reader Adaptation (project page), accepted at NeurIPS 2026, asks how memories built by one model can be read by another. Follow-up work is under review. MemoryAthena (project page) routes adaptively between latent and generated memories, and Label-Free Steering (project page) compresses test-time reinforcement learning into bias-only subspaces. The latter is our first step into multimodality. I hope to continue and expand this line of work in the future, and to see more students and researchers advance through their next milestones.
AI for health and mental health
In mental health, our ACL 2026 Findings paper You Never Know a Person, You Only Know Their Defenses introduced the task of detecting levels of psychological defense mechanisms in supportive conversations. That task became a shared task (more below). SQPsych on generating synthetic therapist-client conversations from questionnaires was accepted at EMNLP 2026 workshops. A reprint on interpretable depression detection is under review. On the health side, PhysAssistBench (EMNLP 2026 Findings) evaluates whether LLMs are ready to assist physicians in interactive doctor, patient and EHR settings.
One more paper that does not fit neatly into these lines: a perspective on LLMs for graph learning, Graph2text or Graph2token, appeared in ACM Transactions on Information Systems.
Of course, some (perhaps many) of our submissions were rejected at venues such as ACL, EMNLP and NeurIPS. This is a normal part of research. I am glad to see that those rejected papers are still read and cited by others, which is a good sign that the work is useful to the community.
Funding and a national health model
The biggest new project this year is the FINe-Health Foundry. It is a consortium, funded by Business Finland's Rise to Challenge (Näytönpaikka) programme, that is building a nationwide foundation health model from Finnish national health data. The consortium is led by ELLIS Institute Finland with Aalto University, the University of Helsinki and the University of Turku. Our group at the University of Turku contributes to three research lines: pretraining the foundation health model, medical reasoning and grounding, and agentic "virtual lab" methods where clinicians can audit and give feedback to medical agents.
On the compute side, Microsoft's LINGUA programme supported FineOPUS with compute credits.
Unfortunately, my application to the Research Council of Finland's Academy Project call was unsuccessful. The main criticism I read in the reviews was that the proposal was too ambitious. Even so, I hope the group will keep working on ambitious problems; that is why we are researchers. I also hope to see more funding opportunities for ambitious research in Finland, and more researchers and students taking on ambitious problems.
Open science
FineOPUS is our attempt to do for parallel data what FineWeb did for web text: take the open OPUS repository and refine it into a clean, high-quality corpus. The current release has 1.65 trillion tokens across 9,532 language pairs, filtered from 83.9 billion raw lines down to 36.1 billion high-quality lines. In June I presented FineOPUS at the LINGUA Community Day at the Council of Europe in Strasbourg, and pitched it at an MEP dinner at the European Parliament. The project is an example of open science, and of ambition, in action. The dataset, the pipeline and a technical report will all be released openly. Best wishes to the project!
We also developed and released GlotEval, a massively multilingual evaluation suite for LLMs, which is open source and available on GitHub. Somewhat disappointingly, the suite is not actively maintained at the moment, but I hope to find funding to continue the work and extend it to more languages and tasks.
I also contributed to the Last Translation Benchmark, a large community effort to collect inputs that current translation systems still get wrong. What I learned from writing items by hand and with an agentic pipeline is in my previous post.
Supervision and teaching
Much of this work belongs to the students and junior researchers I am lucky to work with. I supervise three PhD students and co-supervise two more, at the University of Helsinki and the University of Manchester. The process is sometimes hard: the research itself is difficult, our review system is somewhat broken, and competition in our field is harsh. But we all need to keep calm, carry on, and work on things that are worth doing.
I also supervised four MSc theses at the University of Jyväskylä, the University of Turku and Aalto University. Two are completed, on multilingual NLP and multimodal AI. The first grew into our preprint Model-Based Quality Assessment for Massively Multilingual Parallel Data, and a preprint based on the second is in preparation. Two more, on AI for health and AI agents, are in progress.
In teaching, I co-teach the Deep Learning in Language Technology course at the University of Turku.
Community
Two events I co-organized this year both grew out of our mental health research. The PsyDefDetect Shared Task at BioNLP 2026, co-located with ACL 2026 in San Diego, asked systems to detect levels of psychological defense mechanisms in emotional support conversations (overview paper). It drew 172 registered participants, 563 submissions and 21 ranked teams, well beyond what we expected. The MultiPsyche Workshop at AACL-IJCNLP 2026 is the first workshop on multimodal, multilingual and multicultural mental health and psychotherapy. It takes place on 9 November in Hengqin, Zhuhai.
I also served as Publication Chair of EMNLP 2026 and Workshop Chair of NoDaLiDa 2027. At ELLIS Institute Finland and the University of Turku, I sat on the search committee for the institute's second PI and joined the buddy programme for new PIs. I am also on the secondment committee of the HAIF COFUND programme.
I gave two invited talks: Continual Training, Reasoning, and Evaluation in Multilingual LLMs at the TruX Distinguished Seminar Series, University of Luxembourg (29 June 2026), and Large Language Models for Scalable Mental Health Support at the Planetary Health Symposium in Turku (10 August 2026).
Looking ahead
The coming year will be shaped by the FINe-Health Foundry and by releasing FineOPUS in full. I also want to keep the reasoning and memory work moving and scaling up, and to see the MSc and PhD students through their next milestones. Thank you to all my students, collaborators and colleagues at ELLIS Institute Finland, TurkuNLP, Helsinki-NLP and all collaborating institutions such as the University of Manchester, LMU Munich and Aalto University. If any of this work interests you, whether as a student, collaborator or sponsor, please get in touch.