2606.26091v1 Jun 24, 2026 cs.LG

샘플링된 데모를 사용한 온폴리치 자기 증류 학습은 출력 다양성을 감소시킨다

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity

Mohammad Pezeshki
Mohammad Pezeshki
Citations: 183
h-index: 7
Aaron C. Courville
Aaron C. Courville
Citations: 212
h-index: 8
Andrei Liviu Nicolicioiu
Andrei Liviu Nicolicioiu
Mila
Citations: 127
h-index: 5

온폴리치 자기 증류 학습은 단일 모델을 교사와 학생으로 모두 사용하여 높은 pass@1 정확도를 달성하며, 교사는 올바른 데모에 조건화되어 상세한 토큰 레벨 피드백을 제공합니다. 본 연구에서는 이러한 방식이 숨겨진 비용을 초래할 수 있음을 보여줍니다. 즉, 롤아웃의 다양성이 감소하고 pass@k 곡선이 평탄해집니다 (즉, 더 많은 롤아웃을 생성해도 정확도가 향상되지 않습니다). 이는 샘플링된 데모를 사용한 자기 증류 학습 설계에 내재된 복합적인 편향으로 인해 발생합니다. 교사는 학생의 각 롤아웃에 대해 점수를 매기는데, 이때 교사는 샘플링된 올바른 롤아웃에 조건화되어 모델 자체의 편향을 통해 피드백을 전달합니다. 본 연구는 최적의 자기 증류 학습 정책을 이론적으로 분석하고, 학생의 롤아웃과 컨텍스트로 사용되는 올바른 롤아웃 간의 pointwise conditional mutual information 점수를 통해 기본 분포가 왜곡된다는 것을 보여줍니다. 이상적인 온폴리치 강화학습 (RL)은 동일하게 올바른 롤아웃 간의 확률 비율을 유지하는 반면, 자기 증류 학습은 기존의 확률 격차를 증폭시켜 이미 우세한 모드에 더 많은 확률 질량을 집중시킵니다. 통제된 그래프 경로 탐색 작업 및 과학 질문 답변 벤치마크에서, 자기 증류 모델은 평균 성능 면에서 RL과 동등하거나 능가하지만, 기능적 및 의미적 다양성이 현저히 낮으며, 다양한 전략이 필요한 out-of-distribution 설정에서는 실패하는 경향을 보입니다.

Original Abstract

On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with the teacher conditioned on a correct demonstration to provide dense token-level feedback. We show that this could come at a hidden cost: rollout diversity decreases and pass@k curves flatten (i.e., generating more rollouts fails to improve accuracy). We trace this to compounding biases in the design of self-distillation with sampled demonstrations. The teacher scores each student rollout while conditioned on a sampled correct rollout, channeling its feedback through the model's own biases. We theoretically analyze the optimal self-distillation policy and show that it tilts the base distribution by a pointwise conditional mutual information score between the student's rollout and the correct rollout used as context. Unlike the ideal optimal on-policy reinforcement learning (RL), which preserves probability ratios among equally correct rollouts, self-distillation can amplify existing probability gaps, concentrating mass on already-dominant modes. On a controlled graph path-finding task and science question-answering benchmarks, self-distilled models match or exceed RL on average performance but exhibit substantially lower functional and semantic diversity, failing on out-of-distribution settings that require diverse strategies.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!