2607.18618v1 Jul 21, 2026 cs.CL

LatentMT: 잠재적 추론을 활용한 기계 번역

LatentMT: Machine Translation with Latent Reasoning

Samar M. Magdy
Samar M. Magdy
The University Of British Columbia
Citations: 116
h-index: 5
Muhammad Abdul-Mageed
Muhammad Abdul-Mageed
Citations: 169
h-index: 8
Wenhui Zhu
Wenhui Zhu
Citations: 10
h-index: 2
Wei-Rui Chen
Wei-Rui Chen
Citations: 71
h-index: 4
Chiyu Zhang
Chiyu Zhang
Meta
Citations: 849
h-index: 13
Zhipeng Wang
Zhipeng Wang
Citations: 26
h-index: 3

본 논문에서는 기계 번역(MT)을 위한 새로운 확장 방법을 제시합니다. 기존 방식이 파라미터 수를 늘리거나 명시적인 사고 과정을 표현하는 토큰을 사용하는 반면, 저희는 추가적인 순환 계산을 사용하여 숨겨진 상태 내에서 잠재적 추론을 활용합니다. Latent-reasoning 루프 언어 모델(LoopLM)을 체계적으로 연구한 첫 번째 사례인 LatentMT를 소개합니다. LatentMT는 가벼운 학습 방식으로 26억 개의 파라미터를 가진 작은 기반 모델을 사용합니다. 고, 중, 저 자원 언어를 포함하는 32개의 번역 방향에서 LatentMT는 크기가 3~5배 더 큰 모델과 동등한 성능을 달성했습니다. LatentMT는 고자원 언어에서 경쟁력을 보이며, 중간 및 저자원 언어에서는 최고 수준의 성능을 보여줍니다. 순환 추론 단계 수를 늘리는 것의 영향을 분석한 결과, 초기에 순환 계산이 번역 품질을 꾸준히 향상시키는 것을 확인했지만, 이후에는 빠르게 포화되는 경향을 보였습니다. 메커니즘적 분석 결과, 순환 추론 단계를 거치면서 숨겨진 표현 간의 차이가 줄어드는 것으로 나타났으며, 이는 성능 포화 현상을 뒷받침합니다. 또한 효율성 분석 결과, LatentMT는 유사한 성능을 보이는 훨씬 큰 비-잠재적 추론 모델보다 학습 및 추론에 필요한 계산량이 적습니다. 따라서 잠재적 순환 계산은 작고 효율적이면서 강력한 기계 번역 시스템을 개발하는 데 유망한 접근 방식입니다.

Original Abstract

Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parameter count or emitting explicit chain-of-thought tokens, they spend additional recurrent computation inside hidden states. We introduce LatentMT, the first systematic study of latent-reasoning LoopLMs for machine translation. LatentMT adapts a small 2.6B-parameter backbone model with lightweight training. Across 32 translation directions spanning high-, mid-, and low-resource languages, LatentMT achieves performance comparable to models three to five times larger. It is competitive in a high-resource language and achieves state-of-the-art performance on both mid-resource and low-resource languages. Studying the behavior of scaling the number of recurrent reasoning steps, we find that recurrent computation consistently improves translation quality in early steps, then saturates quickly afterwards. Our mechanistic analysis shows that hidden-representation differences shrink along the recurrent reasoning-step axis, supporting the observed saturation in performance. Finally, our efficiency analysis shows that LatentMT requires lower training and inference compute than much larger non-latent-reasoning models with similar performance, making latent recurrent computation a promising path toward compact, efficient, and strong machine translation.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!