P3B3: LLM의 유럽 및 브라질 포르투갈어 방언 편향을 측정하기 위한 다중 회화 벤치마크
P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs
대규모 언어 모델(LLM)이 일상적인 의사소통에 점점 더 많이 활용됨에 따라, 신뢰하고 공정한 언어 사용을 위해서는 지역별 언어적 다양성을 포착하는 것이 필수적입니다. 포르투갈어의 경우, 유럽(pt-PT) 및 브라질(pt-BR) 방언은 균등하게 대표되지 않으며, 데이터 양에서는 pt-BR이 압도적으로 많지만, LLM이 포르투갈어 방언을 선호하는 정도는 아직 충분히 연구되지 않았습니다. 이러한 간극을 해소하기 위해, 우리는 전문가가 큐레이팅한 언어 방언에 독립적인 회화 프롬프트 벤치마크인 P3B3를 소개하며, 이를 통해 방언 편향과 제어 가능성을 측정할 수 있는 평가 프레임워크도 함께 제공합니다. 여러 모델에 대한 실험 결과, 대부분의 LLM이 pt-BR에 강한 편향을 보이는 경향이 있으며, 모델별로 제어 가능성에 차이가 있음을 확인했습니다. 이러한 결과는 언어 방언 간 균형 잡힌 다국어 표현의 필요성을 강조합니다.
As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use. In Portuguese, European (pt-PT) and Brazilian (pt-BR) varieties remain unevenly represented, with pt-BR dominating in data quantity, while LLM preference for Portuguese variants remains underexplored. To address this gap, we introduce P3B3, an expert-curated language variety agnostic benchmark of conversational prompts, along with an evaluation framework for measuring variety bias and controllability. Experiments on several models show that most LLMs exhibit a strong bias toward pt-BR, with variation in controllability across models. These results highlight the need for more balanced multilingual representation across language varieties.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.