LQM: 언어학적 기반의 다차원 품질 지표: 기계 번역을 위한 방법
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
기존의 기계 번역 평가 프레임워크, 자동 평가 지표 및 다차원 품질 지표(MQM)와 같은 인간 평가 방식은 대부분 언어에 구애받지 않도록 설계되었습니다. 그러나 이러한 방식은 종종 방언 및 문화적 특성을 고려하지 못하여, 특히 아랍어와 같이 언어 변종, 내용 범위, 그리고 어휘적 적절성 간의 불일치로 인해 번역 오류가 발생하는 이중 언어 환경에서 한계를 드러냅니다. 본 논문에서는 LQM(Linguistically Motivated Multidimensional Quality Metrics)을 소개합니다. LQM은 기계 번역 오류를 진단하기 위한 계층적 오류 분류 체계로, 사회언어학, 화용론, 의미론, 형태구문, 철자법, 그리고 문자 체계의 여섯 가지 언어학적 기반 수준으로 구성됩니다(그림 1). 본 연구에서는 7개의 아랍어 방언(이집트, 아랍에미리트, 요르단, 모리타니, 모로코, 팔레스타인, 예멘)에 걸쳐 대화체 및 문화적으로 풍부한 콘텐츠에서 추출한 3,850개의 문장(각 방언당 550개)으로 구성된 양방향 병렬 코퍼스를 구축했습니다. 우리는 6개의 LLM을 제로샷 환경에서 평가하고, LQM을 사용하여 전문가가 문장 수준으로 인간 주석을 수행하여 3,495개의 고유한 오류 문장에 걸쳐 6,113개의 레이블이 지정된 오류 구간을 생성하고, 심각도에 가중치를 부여한 품질 점수를 산출했습니다. 또한, 자동 평가 지표인 spBLEU를 사용하여 분석을 보완했습니다. 본 연구에서는 아랍어를 대상으로 검증했지만, LQM은 언어에 구애받지 않는 프레임워크로 설계되었으며, 다른 언어에 쉽게 적용하거나 조정할 수 있습니다. LQM에 의해 주석 처리된 오류 데이터, 프롬프트 및 주석 지침은 https://github.com/UBC-NLP/LQM_MT 에서 공개적으로 이용할 수 있습니다.
Existing MT evaluation frameworks, including automatic metrics and human evaluation schemes such as Multidimensional Quality Metrics (MQM), are largely language-agnostic. However, they often fail to capture dialect- and culture-specific errors in diglossic languages (e.g., Arabic), where translation failures stem from mismatches in language variety, content coverage, and pragmatic appropriateness rather than surface form alone.We introduce LQM: Linguistically Motivated Multidimensional Quality Metrics for MT. LQM is a hierarchical error taxonomy for diagnosing MT errors through six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics (Figure 1). We construct a bidirectional parallel corpus of 3,850 sentences (550 per variety) spanning seven Arabic dialects (Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni), derived from conversational, culturally rich content. We evaluate six LLMs in a zero-shot setting and conduct expert span-level human annotation using LQM, producing 6,113 labeled error spans across 3,495 unique erroneous sentences, along with severity-weighted quality scores. We complement this analysis with an automatic metric (spBLEU). Though validated here on Arabic, LQM is a language-agnostic framework designed to be easily applied to or adapted for other languages. LQM annotated errors data, prompts, and annotation guidelines are publicly available at https://github.com/UBC-NLP/LQM_MT.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.