저자원 환경에서의 기계 번역 품질 평가를 위한 도메인 특화 방법
Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios
품질 평가(QE)는 참조 자료가 없는 환경, 특히 도메인 특화 및 저자원 언어 환경에서 기계 번역 품질을 평가하는 데 필수적입니다. 본 논문에서는 영어에서 인도어 기계 번역에 대한 문장 수준의 품질 평가를 수행하며, 의료, 법률, 관광, 일반 분야의 네 가지 도메인과 다섯 가지 언어 쌍을 대상으로 합니다. 우리는 선택된 폐쇄형 및 개방형 LLM에서 제로샷, 소수 샘플, 그리고 지침 기반 프롬프팅을 체계적으로 비교합니다. 연구 결과는 폐쇄형 모델이 프롬프팅만으로도 강력한 성능을 보이지만, 특히 고위험 도메인에서 개방형 모델의 경우 프롬프팅만으로는 여전히 불안정하다는 것을 보여줍니다. 이를 해결하기 위해, 우리는 LLM 기반 품질 평가를 위한 프레임워크인 ALOPE를 사용합니다. ALOPE는 선택된 중간 Transformer 레이어에 회귀 헤드를 연결한 Low-Rank Adaptation을 활용합니다. 또한, 우리는 최근 제안된 Low-Rank Multiplicative Adaptation (LoRMA)을 ALOPE에 통합했습니다. 실험 결과는 중간 레이어 적응이 품질 평가 성능을 지속적으로 향상시키며, 특히 의미적으로 복잡한 도메인에서 성능 향상을 보여주어, 실제 환경에서 더 안정적인 품질 평가를 위한 방법을 제시합니다. 본 연구에서 사용된 코드와 도메인 특화 품질 평가 데이터셋을 공개하여 추가 연구를 지원합니다.
Quality Estimation (QE) is essential for assessing machine translation quality in reference-less settings, particularly for domain-specific and low-resource language scenarios. In this paper, we investigate sentence-level QE for English to Indic machine translation across four domains (Healthcare, Legal, Tourism, and General) and five language pairs. We systematically compare zero-shot, few-shot, and guideline-anchored prompting across selected closed-weight and open-weight LLMs. Findings indicate that while closed-weight models achieve strong performance via prompting alone, prompt-only approaches remain fragile for open-weight models, especially in high-risk domains. To address this, we adopt ALOPE, a framework for LLM-based QE that uses Low-Rank Adaptation with regression heads attached to selected intermediate Transformer layers. We also extend ALOPE with recently proposed Low-Rank Multiplicative Adaptation (LoRMA). Our results show that intermediate-layer adaptation consistently improves QE performance, with gains in semantically complex domains, indicating a path toward more robust QE in practical scenarios. We release code and domain-specific QE datasets publicly to support further research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.