While the advent of large language models (LLMs) has marked a transformative phase in AI, existing models often fall short of the specific needs of the public sector and other users in Europe.
Proprietary and 3rd-party LLMs offer powerful capabilities, but have limitations around language diversity, particularly for low-resource languages (those with fewer speakers and less linguistic data available).
The exact data sources are rarely reported in detail, but a widely used resource is Common Crawl, where most EU languages are highly underrepresented. For example, Latvian accounts for only 0.09% of the total dataset, Irish 0.07% and Maltese 0.03%. The least-represented half of all EU official languages add up to merely 2.4%. Other challenges include data quality, copyright safety, transparency and freedom from bias.
Contributing to the European ecosystem of LLMs
The European Commission is working to address these limitations by using the high-quality multilingual data generated by the EU institutions to contribute to the European ecosystem of LLMs – which are better suited to the EU’s multilingual landscape.
This work is part of DG Translation’s partnership with the Directorate-General for Communications Networks, Content and Technology (DG CONNECT) for AI-based multilingual services under the Digital Europe programme.
AI for a multilingual Europe
Models that cover only a limited number of languages and underperform on low-resource languages are major obstacles for multilingual organisations and societies. This is particularly relevant for a multilingual Europe, especially when it comes to European AI projects that require a broad range of EU languages.
The EU Institutional LLM – enhanced with formal texts from the EU institutions – is a powerful addition to European sovereign AI, complementing other efforts and contributing to a diverse landscape of European solutions. Its enhanced EU language capabilities and EU knowledge are better tailored for use by EU public administrations, small businesses, academia and non-governmental organisations.
The project is a stepping stone in the EU’s ambitions to become a major player in AI innovation and strategic technologies, while being able to rely on its own digital systems and tools.
Creating a high-quality EU Institutional LLM
To realise this vision, DG Translation’s expert engineers are building on an existing European open LLM, with the goal of improving its multilingual capabilities. This work draws on 2 key European assets:
- the supercomputers provided by the European High Performance Computing Joint Undertaking (EuroHPC JU)
- the datasets of the European Advanced Multilingual Information System (Euramis), a unique and voluminous corpus of multilingual text from all the EU institutions.
Better coverage of EU languages
In addition, DG Translation’s proximity to language professionals and their direct feedback gives us a major advantage when it comes to preparing data and evaluating models. The resulting models demonstrate better coverage of all EU official languages and an improved ability to handle EU topics.
Early results confirm this potential. In a benchmark based on EU institutional texts, the EU Institutional LLM consistently outperformed the original model across all EU languages tested. The gains were most significant for languages typically underrepresented in global AI training data:
- Irish nearly quadrupled its score
- Estonian improved by almost 80%
- Greek nearly doubled its result
- Latvian and Lithuanian also recorded gains of around 70 to 75%.
In various stages of the EU Institutional LLM project, the European Commission has been able to access supercomputing infrastructure through 3 projects with the EuroHPC JU. First, the goal was to develop a cutting-edge skillset and demonstrate the ability to train large AI models with the MeluXina supercomputer in Luxembourg. Then followed more advanced and intense training of LLMs with the Leonardo supercomputer in Bologna. Currently, we are testing our algorithm's efficiency on the MareNostrum 5 supercomputer in Barcelona to optimise future training.
The EU Institutional LLM has been built with European technology at its core, which is why an existing European open LLM was selected for the project. Models created by Mistral AI (Mixtral 8x7B and 8x22B) are enhanced with our Euramis language data.
The EU Institutional LLM in practice
The EU Institutional LLM already powers eSummary, one of our AI-based multilingual tools. Future iterations will power more of our tools.
The model is available for download from the European Language Data Space by any EU-based legal entity:
- EU Institutional LLM v1 base – continually pretrained multilingual base model
- EU Institutional LLM v1 instruct – instruction-tuned version of the base model
What are the main features of DG Translation’s project?
- Comprehensive inclusion of all 24 EU official languages
Each low-resource language is represented with at least 1 billion tokens (units of text). - Bottom-up approach
The Euramis data used is meticulously curated in line with stringent quality standards. It is aligned with core European values and devoid of any copyright infringements. - Pre- and post-training
The project focuses on continuing the basic training of a state-of-the-art LLM with EU institutional data, fine-tuning it for use cases in EU-public administrations and aligning it to EU values and EU-specific preferences.
The current model is the first step in a longer-term programme, with future iterations planned.
Measuring multilingual performance: EU MMLU, the EU benchmark for LLMs
Building a model that covers all EU official languages is only part of the challenge. We also need reliable tools to measure how well the model performs across those languages.
LLMs are typically evaluated using benchmarking datasets that contain multiple-choice questions covering different areas of human knowledge. The LLM receives a score based on its answers to these questions.
Most AI benchmarks used today were developed in English and do not specifically reflect European educational, cultural and societal contexts. They do not always give an accurate picture of how a model performs in other languages.
To address this gap, DG Translation has released the EU MMLU, a high-quality multilingual benchmarking dataset that allows AI developers, researchers and public administrations to assess whether LLMs perform fairly and effectively across the EU's linguistic diversity.
It builds on one of the most widely used AI evaluation datasets, Massive Multitask Language Understanding (MMLU). This dataset contains thousands of benchmark questions on 57 subjects, ranging from science, technology and mathematics to law, ethics, economics, religion and public affairs.
Unlike most existing multilingual benchmarks, which rely primarily on machine translation, the EU MMLU dataset uses a human-centred approach. DG Translation has joined forces with student translators and project managers from the European Master's in Translation (EMT) network to translate and revise 1000+ benchmark questions on 7 of these subject areas, selected for their relevance to the EU.
The objective is to make sure the questions retain the same meaning, difficulty and testing value across languages. The dataset is currently available in Croatian, Czech, Dutch, English, French, German, Greek, Hungarian, Irish, Italian, Lithuanian, Polish, Portuguese, Romanian, Slovak and Slovenian. Further languages will be added as the project progresses towards full coverage of all EU official languages.
The EU MMLU can be accessed via
- the European Language Data Space – EU MMLU (open to any EU-based legal entity)
- Hugging Face – EU MMLU (open access).
Quality criteria for EU-oriented multilingual benchmarks
Alongside the EU MMLU dataset, DG Translation has published a list of core quality criteria for EU-oriented multilingual benchmarking of LLMs.
This list is intended to support creators of benchmarking datasets in designing evaluations that work across EU languages and reflect EU values. It can also help builders and providers of LLMs to improve the performance of their models for users in the EU – and to highlight how their models are better suited for EU languages and EU contexts.
Our long-term ambition is to establish a new standard for multilingual AI evaluation in Europe.
This page was last updated on 21 July 2026
