2602.05879v1 Feb 05, 2026 cs.CL

EuroLLM-22B: 기술 보고서

EuroLLM-22B: Technical Report

Nuno M. Guerreiro
Nuno M. Guerreiro
Instituto de Telecomunicações
Citations: 1,990
h-index: 19
José Pombal
José Pombal
Sword Health
Citations: 751
h-index: 10
P. Martins
P. Martins
Citations: 490
h-index: 6
Ricardo Rei
Ricardo Rei
Citations: 430
h-index: 11
Miguel Moura Ramos
Miguel Moura Ramos
Citations: 50
h-index: 4
Duarte M. Alves
Duarte M. Alves
Citations: 1,748
h-index: 14
Hippolyte Gisserot-Boukhlef
Hippolyte Gisserot-Boukhlef
Citations: 194
h-index: 5
João Alves
João Alves
Citations: 586
h-index: 8
Patrick Fernandes
Patrick Fernandes
Citations: 605
h-index: 10
Nicolas Boizard
Nicolas Boizard
Citations: 164
h-index: 4
Amin Farajian
Amin Farajian
Citations: 443
h-index: 6
Mateusz Klimaszewski
Mateusz Klimaszewski
Warsaw University of Technology
Citations: 176
h-index: 6
J. G. D. Souza
J. G. D. Souza
Citations: 33
h-index: 3
Barry Haddow
Barry Haddow
Citations: 161
h-index: 7
Franccois Yvon
Franccois Yvon
Citations: 286
h-index: 7
Pierre Colombo
Pierre Colombo
Citations: 1,199
h-index: 13
Alexandra Birch
Alexandra Birch
Citations: 138
h-index: 4
Andr'e F. T. Martins
Andr'e F. T. Martins
Citations: 292
h-index: 10

본 보고서는 유럽 시민들의 요구를 충족시키기 위해 만들어진 대규모 언어 모델인 EuroLLM-22B를 소개합니다. EuroLLM-22B는 유럽 연합의 24개 공식 언어와 11개의 추가 언어를 지원합니다. 본 모델은 기존의 공개된 대규모 언어 모델에서 유럽 언어가 과소 대표되고 부족하게 제공되는 문제를 해결하고자 합니다. EuroLLM-22B의 개발 과정에 대한 종합적인 개요를 제공하며, 여기에는 토크나이저 설계, 아키텍처 사양, 데이터 필터링 및 학습 절차가 포함됩니다. 다양한 다국어 벤치마크에서 EuroLLM-22B는 추론, 지시 따르기 및 번역 분야에서 뛰어난 성능을 보이며, 유사한 크기의 모델과 경쟁력 있는 결과를 달성했습니다. 향후 연구를 지원하기 위해, 우리는 기본 모델 및 지시-튜닝 모델, 다국어 웹 사전 학습 데이터 및 업데이트된 EuroBlocks 지시 데이터셋, 그리고 사전 학습 및 평가 코드를 공개합니다.

Original Abstract

This report presents EuroLLM-22B, a large language model trained from scratch to support the needs of European citizens by covering all 24 official European Union languages and 11 additional languages. EuroLLM addresses the issue of European languages being underrepresented and underserved in existing open large language models. We provide a comprehensive overview of EuroLLM-22B's development, including tokenizer design, architectural specifications, data filtering, and training procedures. Across a broad set of multilingual benchmarks, EuroLLM-22B demonstrates strong performance in reasoning, instruction following, and translation, achieving results competitive with models of comparable size. To support future research, we release our base and instruction-tuned models, our multilingual web pretraining data and updated EuroBlocks instruction datasets, as well as our pre-training and evaluation codebases.

2 Citations
0 Influential
9.5 Altmetric
49.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!