AMALIA 기술 보고서: 유럽 포르투갈어에 특화된 완전 오픈 소스 대규모 언어 모델
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
오픈 소스 대규모 언어 모델(LLM)의 빠른 발전에도 불구하고, 유럽 포르투갈어(pt-PT)는 학습 데이터와 자체 평가 모두에서 상대적으로 부족한 상황이며, 기계 번역을 통해 만들어진 벤치마크는 해당 언어의 언어적, 문화적 뉘앙스를 제대로 반영하기 어렵습니다. 본 논문에서는 유럽 포르투갈어(pt-PT)에 중점을 둔 완전 오픈 소스 LLM인 AMALIA를 소개합니다. AMALIA는 학습 과정 전반에 걸쳐 고품질의 pt-PT 데이터를 더 많이 활용하여 개발되었습니다. 유럽 포르투갈어를 더욱 정확하게 평가하기 위해, 표준 작업의 번역본과 함께 유럽 포르투갈어 생성, 언어 능력, 그리고 유럽 포르투갈어/브라질 포르투갈어 간의 편향성을 측정하는 네 개의 새로운 데이터셋을 포함하는 pt-PT 벤치마크 모음을 공개합니다. 실험 결과, AMALIA는 번역된 벤치마크에서 강력한 기준 모델과 유사한 성능을 보였으며, 유럽 포르투갈어에 특화된 평가에서는 성능이 크게 향상되었습니다. 이는 유럽 포르투갈어에 대한 맞춤형 학습과 자체 벤치마킹의 중요성을 뒷받침합니다.
Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.