2603.26511v1 Mar 27, 2026 cs.CL

AMALIA 기술 보고서: 유럽 포르투갈어에 특화된 완전 오픈 소스 대규모 언어 모델

AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese

Miguel Moura Ramos
Miguel Moura Ramos
Citations: 50
h-index: 4
Duarte M. Alves
Duarte M. Alves
Citations: 1,748
h-index: 14
Inês Vieira
Inês Vieira
Citations: 21
h-index: 2
I. Calvo
I. Calvo
Citations: 11
h-index: 1
Iago Paulo
Iago Paulo
Citations: 1
h-index: 1
James Furtado
James Furtado
Citations: 0
h-index: 0
Rafael Ferreira
Rafael Ferreira
Nova School of Science and Technology
Citations: 63
h-index: 5
Diogo Tavares
Diogo Tavares
Citations: 52
h-index: 4
Diogo Gl'oria-Silva
Diogo Gl'oria-Silva
Citations: 10
h-index: 1
David Semedo
David Semedo
Universidade NOVA de Lisboa
Citations: 289
h-index: 9
João Magalhães
João Magalhães
Citations: 137
h-index: 5
Afonso Simpl'icio
Afonso Simpl'icio
Citations: 9
h-index: 1
Gonccalo Vinagre
Gonccalo Vinagre
Citations: 0
h-index: 0
Giuseppe Attanasio
Giuseppe Attanasio
Citations: 9
h-index: 2
Rui Guerra
Rui Guerra
Citations: 19
h-index: 3
Beatriz Canaverde
Beatriz Canaverde
Citations: 8
h-index: 2
Vasco Ramos
Vasco Ramos
Citations: 23
h-index: 4
Miguel Faria
Miguel Faria
Citations: 442
h-index: 3
Marcos V. Treviso
Marcos V. Treviso
Citations: 60
h-index: 4
D. Gomes
D. Gomes
Citations: 0
h-index: 0
P. Gomes
P. Gomes
Citations: 0
h-index: 0
André Martins
André Martins
Citations: 1,421
h-index: 13

오픈 소스 대규모 언어 모델(LLM)의 빠른 발전에도 불구하고, 유럽 포르투갈어(pt-PT)는 학습 데이터와 자체 평가 모두에서 상대적으로 부족한 상황이며, 기계 번역을 통해 만들어진 벤치마크는 해당 언어의 언어적, 문화적 뉘앙스를 제대로 반영하기 어렵습니다. 본 논문에서는 유럽 포르투갈어(pt-PT)에 중점을 둔 완전 오픈 소스 LLM인 AMALIA를 소개합니다. AMALIA는 학습 과정 전반에 걸쳐 고품질의 pt-PT 데이터를 더 많이 활용하여 개발되었습니다. 유럽 포르투갈어를 더욱 정확하게 평가하기 위해, 표준 작업의 번역본과 함께 유럽 포르투갈어 생성, 언어 능력, 그리고 유럽 포르투갈어/브라질 포르투갈어 간의 편향성을 측정하는 네 개의 새로운 데이터셋을 포함하는 pt-PT 벤치마크 모음을 공개합니다. 실험 결과, AMALIA는 번역된 벤치마크에서 강력한 기준 모델과 유사한 성능을 보였으며, 유럽 포르투갈어에 특화된 평가에서는 성능이 크게 향상되었습니다. 이는 유럽 포르투갈어에 대한 맞춤형 학습과 자체 벤치마킹의 중요성을 뒷받침합니다.

Original Abstract

Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!