평가 지표는 고전 중국어에서 영어로 번역된 결과물의 오류를 감지할 수 있는가?
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
최근 거대 언어 모델은 일부 역사적 언어를 놀랍도록 잘 번역할 수 있지만, 신뢰성 있는 평가의 부족으로 인해 디지털 인문학 워크플로우에서의 활용성이 제한됩니다. 본 연구에서는 고전 중국어에서 영어로 번역된 결과물을 예시로 사용하여, 현대 언어를 위해 개발된 기존의 자동 평가 지표들이 이러한 맥락에서 얼마나 신뢰할 만한지를 조사합니다. 우리는 학문적 사용에서 중요한 오류 유형을 포착하는 최소 쌍(minimal pairs) 기반의 진단 프레임워크를 도입하여, 참조 기반 및 참조 없는 평가 지표 모두의 오류 감지 능력과 유효한 변동에 대한 허용 범위를 분석합니다. 연구 결과, 모든 평가 지표가 특정 측면에서 한계를 드러냈지만, MetricX24가 전반적으로 가장 우수한 성능을 보였습니다. 이러한 발견은 역사적 및 문화적으로 구별되는 번역 환경을 위한 더욱 강력하고 해석 가능한 평가 지표의 필요성을 강조합니다.
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.