사고를 담은 번역: 강화 학습을 이용한 난이도 적응 추론 - 다중 도메인 기계 번역
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
다중 도메인 기계 번역(MDMT)은 각 도메인의 언어적 복잡성 차이로 인해 독특한 어려움을 야기합니다. 인간 번역가가 난이도에 따라 추론 노력을 조절하는 능력에서 영감을 받아, 본 논문에서는 자원 효율적인 프레임워크인 TwT(Translation with Thought)를 제안합니다. TwT는 직관적 추론과 의도적인 추론 간의 추론 과정을 학습하여 조정합니다. TwT는 두 단계로 훈련됩니다: (1) DeepSeek-R1에서 추출하고 GPT-4o가 인간과 유사한 사고 방식을 반영하도록 재작성한 난이도 정보를 포함하는 긴 추론 과정(chain-of-thought)을 기반으로 한 지도 학습 미세 조정, 그리고 (2) 번역 품질 및 추론 효율성을 최적화하기 위한 하이브리드 보상을 사용한 강화 학습. 본 논문에서는 15개의 다양한 도메인 및 언어 환경(3개 기존 언어 및 59개 새로운 언어 포함)에 대한 벤치마크를 사용하여 세 가지 기본 모델에 대한 실험을 진행했습니다. 그 결과, TwT-7B와 TwT-14B는 훨씬 더 큰 최첨단 추론 모델보다 번역 품질 측면에서 우수한 성능을 보였으며, 토큰 사용량을 32~60% 줄였습니다. 이러한 결과는 번역 동작을 인지적 원칙에 맞추면 강력한 일반화 능력, 높은 번역 품질 및 효율적인 추론이 가능함을 보여줍니다.
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators' ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-thought traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32--60\%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.