2608.08283v2 Aug 08, 2026 cs.CL

평가 지표는 고전 중국어에서 영어로 번역된 결과물의 오류를 감지할 수 있는가?

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

Marine Carpuat
Marine Carpuat
University of Maryland
Citations: 7,483
h-index: 36
O. Quinjica
O. Quinjica
Citations: 2
h-index: 1
Eric Bennett
Eric Bennett
Citations: 40
h-index: 1
Xincheng Yang
Xincheng Yang
Citations: 20
h-index: 3
Andrew D. Schonebaum
Andrew D. Schonebaum
Citations: 152
h-index: 7

최근 거대 언어 모델은 일부 역사적 언어를 놀랍도록 잘 번역할 수 있지만, 신뢰성 있는 평가의 부족으로 인해 디지털 인문학 워크플로우에서의 활용성이 제한됩니다. 본 연구에서는 고전 중국어에서 영어로 번역된 결과물을 예시로 사용하여, 현대 언어를 위해 개발된 기존의 자동 평가 지표들이 이러한 맥락에서 얼마나 신뢰할 만한지를 조사합니다. 우리는 학문적 사용에서 중요한 오류 유형을 포착하는 최소 쌍(minimal pairs) 기반의 진단 프레임워크를 도입하여, 참조 기반 및 참조 없는 평가 지표 모두의 오류 감지 능력과 유효한 변동에 대한 허용 범위를 분석합니다. 연구 결과, 모든 평가 지표가 특정 측면에서 한계를 드러냈지만, MetricX24가 전반적으로 가장 우수한 성능을 보였습니다. 이러한 발견은 역사적 및 문화적으로 구별되는 번역 환경을 위한 더욱 강력하고 해석 가능한 평가 지표의 필요성을 강조합니다.

Original Abstract

Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.

0 Citations
0 Influential
18 Altmetric
90.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!