2603.26516v1 Mar 27, 2026 cs.CL

ALBA: 생성형 LLM의 언어 및 언어학적 측면 평가를 위한 유럽 포르투갈어 벤치마크

ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs

Inês Vieira
Inês Vieira
Citations: 21
h-index: 2
I. Calvo
I. Calvo
Citations: 11
h-index: 1
Iago Paulo
Iago Paulo
Citations: 1
h-index: 1
James Furtado
James Furtado
Citations: 0
h-index: 0
Rafael Ferreira
Rafael Ferreira
Nova School of Science and Technology
Citations: 63
h-index: 5
Diogo Tavares
Diogo Tavares
Citations: 52
h-index: 4
Diogo Gl'oria-Silva
Diogo Gl'oria-Silva
Citations: 10
h-index: 1
David Semedo
David Semedo
Universidade NOVA de Lisboa
Citations: 289
h-index: 9
João Magalhães
João Magalhães
Citations: 137
h-index: 5

대규모 언어 모델(LLM)이 다양한 언어 영역으로 확장됨에 따라, 상대적으로 활용이 적은 언어에서의 성능 평가가 더욱 중요해지고 있습니다. 특히 유럽 포르투갈어(pt-PT)는 기존 학습 데이터와 벤치마크가 주로 브라질 포르투갈어(pt-BR)로 구성되어 있어 더욱 큰 영향을 받습니다. 이러한 문제를 해결하기 위해, 우리는 LLM의 pt-PT 언어 능력 평가를 위해 언어학적 기반으로 설계된 벤치마크인 ALBA를 소개합니다. ALBA는 언어 변종, 문화적 의미, 담론 분석, 말장난, 구문, 형태론, 어휘론, 음운론 등 8가지 언어학적 차원을 포괄합니다. ALBA는 언어 전문가에 의해 수동으로 구축되었으며, LLM을 활용한 평가 시스템과 결합되어 pt-PT에서 생성된 언어의 확장 가능한 평가를 가능하게 합니다. 다양한 모델에 대한 실험 결과, 언어학적 차원별 성능의 편차가 나타났으며, 이는 pt-PT 언어 도구 개발을 지원하는 포괄적이고 언어 변종에 민감한 벤치마크의 필요성을 강조합니다.

Original Abstract

As Large Language Models (LLMs) expand across multilingual domains, evaluating their performance in under-represented languages becomes increasingly important. European Portuguese (pt-PT) is particularly affected, as existing training data and benchmarks are mainly in Brazilian Portuguese (pt-BR). To address this, we introduce ALBA, a linguistically grounded benchmark designed from the ground up to assess LLM proficiency in linguistic-related tasks in pt-PT across eight linguistic dimensions, including Language Variety, Culture-bound Semantics, Discourse Analysis, Word Plays, Syntax, Morphology, Lexicology, and Phonetics and Phonology. ALBA is manually constructed by language experts and paired with an LLM-as-a-judge framework for scalable evaluation of pt-PT generated language. Experiments on a diverse set of models reveal performance variability across linguistic dimensions, highlighting the need for comprehensive, variety-sensitive benchmarks that support further development of tools in pt-PT.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!