2608.03077v1 Aug 04, 2026 cs.CL

PAMT: 프로세스 연계 강화 학습을 이용한 다중 도메인 기계 번역

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Biao Fu
Biao Fu
Xiamen University
Citations: 201
h-index: 8
Yongshi Ye
Yongshi Ye
Citations: 17
h-index: 3
Xiaodong Shi
Xiaodong Shi
Citations: 56
h-index: 5

다중 도메인 기계 번역(MDMT)은 단순히 유창한 문장 생성뿐만 아니라, 도메인 식별, 용어 통제, 스타일 적응과 같은 도메인 특화된 번역 결정을 요구합니다. 대규모 추론 모델(LRM)은 중간 번역 단계를 통해 이러한 결정을 명시적으로 드러내지만, 15개 도메인 및 4가지 번역 방향에 대한 분석 결과, 이러한 명시적인 추론은 양날의 검이라는 것을 알 수 있습니다. 즉, 긴 형식의 어려운 번역에서는 성능이 향상되지만, 용어 중심적이거나 스타일 제약이 강한 환경에서는 종종 문제가 발생합니다. 이는 기존 방법들이 최종 출력 또는 대략적인 경로를 최적화하지만, 실제 최종 번역에 어떤 번역 단계가 도움이 되는지 파악하지 못하기 때문입니다. 이러한 문제를 해결하기 위해, 우리는 콜드 스타트 방식의 도메인 인식 Long-CoT 지도 학습과 강화 학습을 결합한 프로세스 연계 훈련 프레임워크인 PAMT를 제안합니다. PAMT는 최종 번역에 대한 시퀀스 레벨 형식 및 결과 보상과 함께, 각 명시적인 번역 단계가 참조 번역의 가능성을 얼마나 증가시키는지 측정하는 단계별 프로세스 보상을 사용합니다. 두 가지 기본 모델에서 PAMT는 기존 모델보다 성능이 향상되고, 평균적으로 기계 번역 전문 기준 모델보다 우수한 성능을 보이며, 다양한 도메인, OOD(Out-of-Distribution), 다국어 환경에서 강력한 LLM/LRM과 경쟁력 있는 결과를 보여줍니다.

Original Abstract

Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!