2608.06347v1 Aug 06, 2026 cs.CL

RP-OPSD: 추론 피벗 가이드 기반 온폴리시 자기 증류를 활용한 다국어 추론 전이

RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

Shujian Huang
Shujian Huang
Citations: 49
h-index: 2
X. Wang
X. Wang
Citations: 5
h-index: 1
Junxiao Liu
Junxiao Liu
Citations: 76
h-index: 3

다국어 추론 전이는 대규모 언어 모델(LLM)의 추론 능력을 고자원 언어를 넘어 확장하는 데 매우 중요합니다. 온폴리시 자기 증류(OPSD) 및 그 변형은 유망한 패러다임으로 부상했으며, 학생 모델이 생성한 결과에 대한 밀집적인 토큰 레벨 감독 신호를 제공하지만, 이들의 목표는 교차 언어 전이에 가장 중요한 추론 신호들을 명시적으로 우선시하지 않습니다. 본 논문에서는 대상 언어의 추론 과정이 표면 텍스트 생성뿐만 아니라 추론 과정을 발전시키거나 방향을 전환하고 이후 추론에 영향을 미치는 '추론 피벗(reasoning pivots)' 생성을 포함한다는 점을 밝힙니다. 따라서 우리는 이러한 피벗 주변에서 집중적인 증류를 수행하는 것이 중요하다고 제안합니다. 이에 따라, 본 논문에서는 영어 참조 솔루션 유무에 따른 매칭된 교사 모델의 분포 변화를 활용하여 집중적인 증류 및 참조 고정을 안내하는 RP-OPSD (Reasoning-Pivot-guided On-Policy Self-Distillation)라는 새로운 방법을 제시합니다. 17개 언어와 다양한 난이도를 포괄하는 수학적 추론 벤치마크 실험 결과, 제안하는 방법은 강력한 다국어 추론 기준 모델 및 OPSD 변형보다 우수한 성능을 보였습니다. 추가 분석 결과, RP-OPSD는 추론 제어 및 문제 조건에 따른 상태 업데이트 토큰에 집중적인 증류를 수행하는 반면, 주로 표면 실현을 지원하는 토큰에는 가중치를 낮추는 것을 확인했습니다. 본 논문의 코드는 다음 주소에서 확인할 수 있습니다: https://github.com/NJUNLP/RP-OPSD.

Original Abstract

Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising paradigm, providing dense token-level supervision on student-generated rollouts, yet their objectives do not explicitly prioritize reasoning signals most critical to cross-lingual transfer. We characterize that target-language reasoning comprises the generation of both surface text and reasoning pivots, which are decisions that advance or redirect the reasoning process and shape subsequent inference. This motivates concentrating privileged distillation around such pivots. We therefore propose RP-OPSD, Reasoning-Pivot-guided On-Policy Self-Distillation, using the distributional shift between matched teacher views with and without an English reference solution as an operational proxy to guide privileged distillation and reference anchoring. Experiments on mathematical reasoning benchmarks covering 17 languages and multiple difficulty levels show that our method outperforms strong multilingual reasoning baselines and OPSD variants. Further analysis reveals that RP-OPSD concentrates privileged distillation on reasoning-control and problem-condistioned state-update tokens, while downweighting it for tokens that mainly support surface realization. Our code is available at https://github.com/NJUNLP/RP-OPSD.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!