EuroLLM-22B: 기술 보고서
EuroLLM-22B: Technical Report
본 보고서는 유럽 시민들의 요구를 충족시키기 위해 만들어진 대규모 언어 모델인 EuroLLM-22B를 소개합니다. EuroLLM-22B는 유럽 연합의 24개 공식 언어와 11개의 추가 언어를 지원합니다. 본 모델은 기존의 공개된 대규모 언어 모델에서 유럽 언어가 과소 대표되고 부족하게 제공되는 문제를 해결하고자 합니다. EuroLLM-22B의 개발 과정에 대한 종합적인 개요를 제공하며, 여기에는 토크나이저 설계, 아키텍처 사양, 데이터 필터링 및 학습 절차가 포함됩니다. 다양한 다국어 벤치마크에서 EuroLLM-22B는 추론, 지시 따르기 및 번역 분야에서 뛰어난 성능을 보이며, 유사한 크기의 모델과 경쟁력 있는 결과를 달성했습니다. 향후 연구를 지원하기 위해, 우리는 기본 모델 및 지시-튜닝 모델, 다국어 웹 사전 학습 데이터 및 업데이트된 EuroBlocks 지시 데이터셋, 그리고 사전 학습 및 평가 코드를 공개합니다.
This report presents EuroLLM-22B, a large language model trained from scratch to support the needs of European citizens by covering all 24 official European Union languages and 11 additional languages. EuroLLM addresses the issue of European languages being underrepresented and underserved in existing open large language models. We provide a comprehensive overview of EuroLLM-22B's development, including tokenizer design, architectural specifications, data filtering, and training procedures. Across a broad set of multilingual benchmarks, EuroLLM-22B demonstrates strong performance in reasoning, instruction following, and translation, achieving results competitive with models of comparable size. To support future research, we release our base and instruction-tuned models, our multilingual web pretraining data and updated EuroBlocks instruction datasets, as well as our pre-training and evaluation codebases.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.