2607.02234v1 Jul 02, 2026 cs.AI

정제된 온폴리시 셀프 증류 (OPSD): 사고 방식을 잃지 않고 수행하는 온폴리시 셀프 증류

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Haobo Wang
Haobo Wang
Citations: 1,249
h-index: 17
Wen-song Ye
Wen-song Ye
Citations: 314
h-index: 7
Junbo Zhao
Junbo Zhao
Citations: 289
h-index: 7
Zhanming Shen
Zhanming Shen
Citations: 87
h-index: 3
Xiaomeng Hu
Xiaomeng Hu
Citations: 182
h-index: 4
Gang Chen
Gang Chen
Citations: 586
h-index: 10
Hao Chen
Hao Chen
Citations: 102
h-index: 4
Shaotian Yan
Shaotian Yan
Citations: 228
h-index: 7
Rui Miao
Rui Miao
Citations: 240
h-index: 5
Chen Shen
Chen Shen
Citations: 25
h-index: 3
Jieping Ye
Jieping Ye
Citations: 553
h-index: 12
Jintao Tong
Jintao Tong
Citations: 93
h-index: 6

온폴리시 셀프 증류(OPSD)는 LLM의 추론 능력을 향상시키는 유망한 방법론으로, 참조 솔루션에 접근할 수 있는 특권적인 '선생 모델'이 학생 모델이 생성한 결과물에 토큰 수준의 감독 신호를 제공합니다. 그러나 본 연구에서는 OPSD가 긴 연쇄적 사고(long-CoT) 추론 모델에서 지속적으로 실패하며, 최대 효과는 미미하고 오히려 이러한 모델들이 의존하는 성찰적 추론 능력을 불안정하게 만드는 것을 발견했습니다. 교사 모델의 감독 신호를 새롭게 분해하여 분석한 결과, 다음과 같은 근본적인 원인을 밝혀냈습니다. 즉, 교사 모델의 감독 신호는 참조 데이터에 의해 유발되는 구성 요소로 인해 지배되며, 이는 참조 데이터에 특화된 단기 기억을 강화하는 경향이 있습니다. 반면, 질문 조건부이며 추론 가능성을 높이는 구성 요소는 무시되거나 적극적으로 억제됩니다. 이러한 분석 결과를 바탕으로, 두 단계로 구성된 해결 방안을 제안합니다. 첫째, 참조 데이터만을 사용하여 학습된 '선생 모델'을 구축하여 감독 신호의 전이 불가능한 구성 요소를 분리합니다. 이렇게 분리된 나머지 부분은 질문 조건부이며 추론 가능성을 높이는 수정 사항을 포함합니다. 둘째, 포인트 뮤추얼 정보(PMI)를 사용하여 이 잔여 부분을 학생 모델이 직접 학습할 수 있는 적절한 PMI 목표 분포로 변환하여 참조 데이터에 의해 유발되는 단기 기억 현상을 제거합니다. 두 개의 데이터셋에서 사용된 네 가지 long-CoT 모델에 대한 실험 결과, 제안하는 방법은 기본 모델과 표준 OPSD 모두보다 일관되게 성능 향상을 보여주었으며, 학습 과정 전반에 걸쳐 모델의 자연스러운 지식적 행동을 유지했습니다.

Original Abstract

On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-level supervision on the student's own generated trajectories. However, we find that OPSD consistently fails on long chain-of-thought (long-CoT) reasoning models, yielding at best marginal gains while destabilizing the reflective reasoning capability these models depend on. Through a novel decomposition of the teacher's supervision signal, we identify the root cause: the teacher's supervision is dominated by a reference-induced component that drives rote memorization of reference-specific shortcuts, while the question-conditioned, inference-transferable component is ignored or actively opposed. Based on this diagnosis, we propose a two-step solution. First, we construct a reference-only teacher (the same model conditioned on the reference without the question) to isolate the non-transferable component of the supervision signal; the residual after subtracting this component captures the question-conditioned, inference-transferable correction. Second, we use pointwise mutual information (PMI) as the mechanism to transform this residual into a well-formed PMI target distribution that the student can directly distill from, filtering out the reference-induced shortcut. Experiments on four long-CoT models across two datasets demonstrate consistent improvements over both the base model and standard OPSD, while preserving the models' natural epistemic behavior throughout training.

6 Citations
1 Influential
8.5 Altmetric
50.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!