2607.02235v1 Jul 02, 2026 cs.CL

다국어 환경 및 저자원 언어에서의 LLM 기반 평가 시스템의 과제와 권장 사항

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

D. Adelani
D. Adelani
Citations: 34
h-index: 4
Jakob Prange
Jakob Prange
Citations: 3
h-index: 1
A. Dougruoz
A. Dougruoz
Citations: 42
h-index: 3
Xixian Liao
Xixian Liao
Citations: 0
h-index: 0
Verena Blaschke
Verena Blaschke
LMU Munich
Citations: 188
h-index: 9
Senyu Li
Senyu Li
Citations: 15
h-index: 2

LLM(Large Language Model) 기반 평가 시스템은 기존 평가 지표의 한계점과 인간 판단과의 높은 상관관계 덕분에 자연어 생성 작업에서 주요 평가 패러다임으로 자리 잡았습니다. 하지만, 대부분 영어에 국한되어 있습니다. 최근에는 다국어 환경, 특히 저자원 언어로 LLM 기반 평가 시스템을 확장하려는 시도가 이루어지고 있습니다. 그러나 LLM은 저자원 언어에 대한 이해 능력이 제한적이며, 이러한 환경에서는 적절한 인간 검증이 부족한 경우가 많습니다. 본 연구는 문제의 범위와 현재 실태를 파악하기 위해 ACL Anthology 논문 중 다국어 환경 및 저자원 언어를 다루는 다양한 작업들을 중심으로 LLM 기반 평가 시스템 사용 사례를 분석했습니다. 650개의 논문 중 LLM 기반 평가 시스템을 언급한 논문은 33개에 불과했으며, 이들 대부분이 저자원 또는 다국어 환경을 다루고 있었습니다. 이러한 논문에 대한 심층적인 분석 결과, 일관성 없는 평가 결과, 다국어 환경에서 LLM 판단을 과도하게 신뢰하는 경향, 그리고 연구 당 단일 모델만을 사용하는 경우가 많다는 점 등이 확인되었습니다. 본 연구는 NLP 커뮤니티에 도움이 되도록 다국어 및 저자원 환경에서 LLM 기반 평가 시스템을 효과적으로 활용하기 위한 권장 사항을 제시합니다.

Original Abstract

LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!