2608.05802v1 Aug 06, 2026 cs.CL

다국어 수학적 추론을 위한 온라인 정책 기반 델타 증류

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo
Byeongho Heo
NAVER AI LAB
Citations: 4,294
h-index: 21
Dongyoon Han
Dongyoon Han
NAVER AI LAB
Citations: 12,036
h-index: 29
Jaehui Hwang
Jaehui Hwang
Citations: 10
h-index: 2
Sangdoo Yun
Sangdoo Yun
Citations: 35
h-index: 4

온라인 정책 기반 증류(OPD)는 LLM의 추가 학습에 있어 강화 학습의 유망한 대안으로 부상하고 있지만, 다국어 환경에서의 효과성은 아직 충분히 연구되지 않았습니다. 본 논문에서는 영어, 한국어 및 일본어로 된 수학적 추론 문제를 해결하기 위해 OPD와 그 발전된 형태인 온라인 정책 기반 델타 증류(OPD$^2$)를 연구합니다. OPD$^2$는 추가 학습된 모델(teacher)과 기본 모델 간의 확률 차이를 학습 신호로 사용하여 OPD를 개선합니다. Qwen3 모델에 대한 실험 결과, OPD$^2$는 원래의 OPD보다 일관되게 우수한 성능을 보이며, 특히 한국어와 일본어에서 상당한 개선이 이루어졌고, 일반적으로 영어-한국어 성능 격차를 줄이는 경향을 나타냅니다. 또한, 영어만으로 학습된 OPD도 한국어 및 일본어 문제 해결 능력을 향상시킬 수 있지만, 종종 응답 방향이 영어로 치우치는 현상이 발생하며, 이는 타겟 언어의 응답을 유지하는 데 있어 다국어 데이터의 중요성을 강조합니다.

Original Abstract

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!