European Project for Linguistic and Technological Sovereignty
Europe is embarking on an ambitious journey toward technological independence in the field of artificial intelligence. At the forefront of this effort is Charles University, which has taken on the role of the main coordinator of the prestigious OpenEuroLLM project. This large-scale project, funded by the European Union under the Horizon Europe programme, aims to create open, multilingual and high-quality next-generation language models for European commercial and public services. The project was officially launched on 1 May 2023 and will run for 36 months, with a budget of nearly 46 million euros, making it one of the most significant European projects in the field of artificial intelligence.
Professor Jan Hajič from the Faculty of Mathematics and Physics at Charles University, who serves as the project's main coordinator, explains the essence of large language models (LLMs) in an interview with Seznam Zprávy: "These are programmes that generate additional words based on the input text, and can gradually create entire sentences or even books. They take linguistic and content knowledge into account, store specific data in their internal tables and continue training on it. The result is usually convincing, but not always factually accurate." This technology, which has made tremendous progress in recent years thanks to models such as ChatGPT, Claude and Gemini, is becoming a strategic technology of the future.
Three Main Pillars of the OpenEuroLLM Project
The OpenEuroLLM project differs from commercial American or Chinese models in three fundamental respects. The first is its emphasis on multilingualism, with high-quality support for all European languages being a top priority. The project is committed to covering all 24 official languages of the European Union, as well as other European languages. While global commercial models such as ChatGPT are also multilingual, the European project places this feature first in order to ensure that smaller language communities also have access to the latest technologies. The second key aspect is openness—the models will be developed as open source and with full transparency, enabling their widespread use across Europe and making it easier to verify compliance with European legislation, particularly the recently approved AI Act. As the project's official website states, all language models created will be available under an open licence, allowing them to be used free of charge for both research and commercial purposes. This makes OpenEuroLLM fundamentally different from closed commercial models, whose users often do not know the exact details of how they function and are trained. The third pillar is the democratisation of access to advanced AI technologies. "We want to minimise digital inequality between languages," Professor Hajič emphasises, adding that the aim is to make cutting-edge technologies accessible not only to large corporations, but also to smaller entities and the public sector. This is in line with European values of inclusivity and equal access to technology.
The Technical Side of the Project and the Consortium of Institutions
The OpenEuroLLM project is being implemented by a consortium of 25 partners from 15 European countries. In addition to Charles University, significant scientific and research institutions such as CNRS (France), Aalto University (Finland), DFKI (Germany), KU Leuven (Belgium), the University of Amsterdam (the Netherlands) and the Barcelona Supercomputing Center (Spain) are involved in the project. This combination of academic institutions, research centres and industrial partners ensures that the project will have access to top-level expertise and infrastructure. From a technical perspective, OpenEuroLLM focuses on developing several types of language models. According to information from the project's official website, both foundation models and specialised models for specific applications will be developed. The models will be released gradually in several versions, with each new version expected to bring improved performance and expanded functionality. Professor Hajič states in the interview: "Our data sets are smaller in scope than those of global players such as OpenAI or Google—but they are of very high quality." This approach reflects the European philosophy of prioritising quality over quantity and ethical considerations over the ruthless maximisation of performance. The models will be trained using the computing capacities of European supercomputers, with access already secured to supercomputers such as LUMI in Finland and MareNostrum in Spain, as stated on the website of the Faculty of Mathematics and Physics at Charles University. This computing infrastructure provides the essential capacity for training large language models, which require enormous computing power.
Real-World Impacts and Practical Applications
The resulting language models are intended to serve a broad range of users—from commercial companies and industrial enterprises to public services throughout Europe. The project's website emphasises that OpenEuroLLM is not merely a research project, but focuses on creating practically usable tools that can be deployed in real-world applications. Potential areas of use include customer support automation, content generation, document translation, text analysis and decision-making support. Information on the project's website indicates that OpenEuroLLM is collaborating with other European AI initiatives, such as the European Language Grid and AI4EU, to ensure compatibility and synergy among various European projects. This cooperation is important for creating a comprehensive European artificial intelligence ecosystem that will be competitive on a global scale.
Timeline and Expected Results
According to information from the project's official sources, the first versions of the models are expected during 2025, while the final models should be completed by the end of the project in April 2026. In the interview, Professor Hajič emphasises that the aim is to create models that will be competitive not only today but also over the next several years, which requires continuous development and improvement. The website of the Faculty of Mathematics and Physics at Charles University states that the project also includes the creation of robust technical infrastructure for training, testing and deploying language models. This infrastructure will remain available to European researchers after the project itself ends, ensuring the long-term sustainability and further development of European language models.
European Digital Sovereignty
The OpenEuroLLM project represents a significant step toward strengthening Europe's competitiveness in artificial intelligence, which is becoming an increasingly important factor in economic growth and social development. As Professor Hajič summarises in his interview with Seznam Zprávy: "Transparent open-source models will strengthen the ability of European companies to compete globally and enable more efficient public services." Information on the project's website indicates that OpenEuroLLM is not merely a technical project, but also has a significant political and strategic dimension. Europe realises that without its own technological capabilities in artificial intelligence, it will become increasingly dependent on technologies developed in the USA or China, which may affect not only the economy, but also security, cultural identity and the ability to promote European values in the digital space. Charles University and the Czech scientist Jan Hajič are thus at the very centre of Europe's efforts toward technological independence in the era of digital transformation. As stated on the website of the Faculty of Mathematics and Physics at Charles University, participation in the OpenEuroLLM project is recognition of the long-standing excellence of Czech research in computational linguistics and artificial intelligence. At the same time, it is an opportunity for Czech researchers to help shape technologies that will significantly influence the lives of millions of Europeans in the years to come.



