다국어 수학적 추론을 위한 온라인 정책 기반 델타 증류
On-Policy Delta Distillation for Multilingual Math Reasoning
온라인 정책 기반 증류(OPD)는 LLM의 추가 학습에 있어 강화 학습의 유망한 대안으로 부상하고 있지만, 다국어 환경에서의 효과성은 아직 충분히 연구되지 않았습니다. 본 논문에서는 영어, 한국어 및 일본어로 된 수학적 추론 문제를 해결하기 위해 OPD와 그 발전된 형태인 온라인 정책 기반 델타 증류(OPD$^2$)를 연구합니다. OPD$^2$는 추가 학습된 모델(teacher)과 기본 모델 간의 확률 차이를 학습 신호로 사용하여 OPD를 개선합니다. Qwen3 모델에 대한 실험 결과, OPD$^2$는 원래의 OPD보다 일관되게 우수한 성능을 보이며, 특히 한국어와 일본어에서 상당한 개선이 이루어졌고, 일반적으로 영어-한국어 성능 격차를 줄이는 경향을 나타냅니다. 또한, 영어만으로 학습된 OPD도 한국어 및 일본어 문제 해결 능력을 향상시킬 수 있지만, 종종 응답 방향이 영어로 치우치는 현상이 발생하며, 이는 타겟 언어의 응답을 유지하는 데 있어 다국어 데이터의 중요성을 강조합니다.
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.