2604.26553v1 Apr 29, 2026 cs.CL

TLPO: 토큰 수준 정책 최적화를 통한 대규모 언어 모델의 언어 혼동 완화

TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models

Naman Goyal
Naman Goyal
Citations: 3,241
h-index: 2
Jerry Tworek
Jerry Tworek
Citations: 47,935
h-index: 13
Mo Bavarian
Mo Bavarian
Citations: 50,094
h-index: 12
Kartikay Khandelwal
Kartikay Khandelwal
Citations: 24,649
h-index: 6
Reiichiro Nakano
Reiichiro Nakano
Citations: 36,785
h-index: 10
Christopher Hesse
Christopher Hesse
Citations: 74,856
h-index: 9
Vineet Kosaraju
Vineet Kosaraju
Citations: 18,140
h-index: 12
K. Cobbe
K. Cobbe
Citations: 16,732
h-index: 11
Dejian Yang
Dejian Yang
Citations: 12,758
h-index: 11
Haowei Zhang
Haowei Zhang
Citations: 10,217
h-index: 8
Jun-Mei Song
Jun-Mei Song
Citations: 17,942
h-index: 11
Qihao Zhu
Qihao Zhu
Citations: 19,846
h-index: 13
Daya Guo
Daya Guo
Citations: 1,063
h-index: 6
Vishrav Chaudhary
Vishrav Chaudhary
Citations: 21,303
h-index: 32
Guillaume Wenzek
Guillaume Wenzek
Citations: 14,486
h-index: 17
Lukasz Kaiser
Lukasz Kaiser
Citations: 30,500
h-index: 8
Albert Q. Jiang
Albert Q. Jiang
Citations: 2,796
h-index: 11
Alexandre Sablayrolles
Alexandre Sablayrolles
Citations: 6,054
h-index: 13
Dan Hendrycks
Dan Hendrycks
Citations: 1,925
h-index: 17
Jinho Choo
Jinho Choo
Citations: 936
h-index: 4
Junseung Lee
Junseung Lee
Citations: 4
h-index: 1
Jimyeong Kim
Jimyeong Kim
Citations: 39
h-index: 4
Y. Song
Y. Song
Citations: 0
h-index: 0
S. K. Hong
S. K. Hong
Citations: 55
h-index: 3
Yeong-Dae Kwon
Yeong-Dae Kwon
Citations: 8
h-index: 2
Peter Clark
Peter Clark
Citations: 0
h-index: 0
Isaac Cowhey
Isaac Cowhey
Citations: 5,056
h-index: 4
Oren Etzioni
Oren Etzioni
Citations: 941
h-index: 9
Tushar Khot
Tushar Khot
Allen Institute for Artificial Intelligence
Citations: 21,752
h-index: 44
Mark Chen
Mark Chen
Citations: 1,053
h-index: 8
Matthias Plappert
Matthias Plappert
Citations: 7
h-index: 2
Jacob Hilton
Jacob Hilton
Citations: 409
h-index: 5
Alexis Conneau
Alexis Conneau
Citations: 4,380
h-index: 5
Peiyi Wang
Peiyi Wang
Citations: 346
h-index: 4
Runxin Xu
Runxin Xu
Citations: 1,803
h-index: 15
Ruoyu Zhang
Ruoyu Zhang
Citations: 82
h-index: 1
Collin Burns
Collin Burns
Anthropic
Citations: 17,769
h-index: 9
Steven Basart
Steven Basart
University of Chicago
Citations: 24,852
h-index: 16
Andy Zou
Andy Zou
Citations: 894
h-index: 10
Antoine Roux
Antoine Roux
Citations: 2,004
h-index: 4
J. Kirkpatrick
J. Kirkpatrick
Citations: 18,249
h-index: 14
Razvan Pascanu
Razvan Pascanu
Citations: 3
h-index: 1
Neil C. Rabinowitz
Neil C. Rabinowitz
Citations: 17,848
h-index: 23
J. Veness
J. Veness
Citations: 48,813
h-index: 24
Guillaume Desjardins
Guillaume Desjardins
Citations: 3,268
h-index: 3
Andrei A. Rusu
Andrei A. Rusu
Citations: 55,201
h-index: 21
Kieran Milan
Kieran Milan
Citations: 13,582
h-index: 6
John Quan
John Quan
Citations: 19,472
h-index: 16
Tiago Ramalho
Tiago Ramalho
Citations: 12,928
h-index: 10

대규모 언어 모델(LLM)은 뛰어난 다국어 능력을 보여주지만, 종종 의도된 언어로 일관된 응답을 생성하지 못하고, '언어 혼동'이라는 현상을 나타냅니다. 기존의 시퀀스 수준 미세 조정 기반 접근 방식(예: DPO, ORPO, GRPO)은 전체 응답 수준에서 작동하며, 모델의 일반적인 능력이 의도치 않게 저하될 수 있다는 문제가 있습니다. 이러한 문제를 해결하기 위해, 우리는 언어 혼동을 완화하기 위한 토큰 수준 업데이트를 기반으로 하는 미세 조정 프레임워크인 '토큰 수준 정책 최적화(TLPO)'를 제안합니다. TLPO는 오류 발생 가능성이 높은 위치를 식별하고, 대체 후보 토큰을 탐색하며, 맞춤형 목표 함수를 사용하여 오류를 유발하는 출력을 세밀한 수준에서 억제하도록 정책을 업데이트합니다. 이러한 선택적인 개입을 통해, 모델의 일반적인 능력에 영향을 주지 않으면서 언어 혼동을 효과적으로 완화할 수 있습니다. 다양한 언어 모델에 대한 실험 결과, TLPO는 언어 일관성을 향상시키는 동시에 하위 작업의 정확도를 유지하면서 기존 방식보다 훨씬 우수한 성능을 보였습니다.

Original Abstract

Large language models (LLMs) demonstrate strong multilingual capabilities, yet often fail to consistently generate responses in the intended language, exhibiting a phenomenon known as language confusion. Prior mitigation approaches based on sequence-level fine-tuning, such as DPO, ORPO, and GRPO, operate at the level of entire responses and can lead to unintended degradation of general model capabilities, motivating the need for more fine-grained alternatives. To address this, we introduce Token-Level Policy Optimization (TLPO), a fine-tuning framework designed to mitigate language confusion through localized, token-level updates. TLPO identifies error-prone positions, explores alternative candidate tokens, and updates the policy using a tailored objective to suppress error-inducing outputs at a granular level. This selective intervention enables effective mitigation of language confusion without compromising the model's general abilities. Experiments on multiple multilingual LLMs across diverse languages demonstrate that TLPO significantly outperforms baselines in improving language consistency while preserving downstream task accuracy.

0 Citations
0 Influential
22 Altmetric
110.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!