대조적 오류 범위 주석(cESA): 여러 번역 결과에 대한 인간 평가
Contrastive ESA: Human Evaluation of Multiple Translations at Once
현재 기계 번역의 인간 평가는 일반적으로 개별 출력 결과를 독립적으로 평가하는 방식으로 진행되며, 이 방식은 높은 평가자 간 불일치와 비용 문제를 야기합니다. 본 연구에서는 원본 입력(텍스트, 비디오, 오디오, 이미지)에 대한 여러 번역 결과를 동시에 제시하는 방법인 대조적 오류 범위 주석(Contrastive Error Span Annotation, cESA) 프로토콜을 소개합니다. cESA에서 평가자는 동일 문서의 여러 번역 결과를 보고 주요 및 경미한 오류 범위를 표시하고, 0%부터 100%까지 절대적인 기준으로 점수를 부여합니다. cESA는 평가자가 여러 출력 결과 간의 공유된 맥락에 접근할 수 있도록 함으로써 더욱 일관성 있고 효율적인 판단을 가능하게 합니다. 본 연구에서는 12개 모델의 영어-일본어 번역에 대한 대규모 인간 평가를 통해 cESA를 검증하고, 표준 점수 기반 평가 방식과 비교하여 주석 시간 단축 및 불일치 감소 효과를 확인했습니다. 기존의 대조 순위 결정 방법과는 달리, cESA는 절대적인 품질 판단 결과를 제공하여, 사후 수정 없이 간단하고 해석 가능한 비모수 모델 순위를 매길 수 있도록 합니다.
Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.