2608.06243v1 Aug 06, 2026 cs.AI

DASH: 추론 모델의 온-폴리시 자체 증류를 위한 발산 적응형 감독 범위

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

Xinyu Tang
Xinyu Tang
Citations: 4,823
h-index: 8
Yafeng Deng
Yafeng Deng
Citations: 35
h-index: 3
Yunyun Han
Yunyun Han
Citations: 11
h-index: 2
Zhiyan Hou
Zhiyan Hou
Citations: 2
h-index: 1
Haiyun Guo
Haiyun Guo
Citations: 1,713
h-index: 20
Jinqiao Wang
Jinqiao Wang
Citations: 49
h-index: 3
Gengsheng Li
Gengsheng Li
Citations: 36
h-index: 2
Xiangzhao Hao
Xiangzhao Hao
Citations: 34
h-index: 4
Hongyan An
Hongyan An
Citations: 6
h-index: 2
Jianjin Zhang
Jianjin Zhang
Citations: 1,257
h-index: 11
Wenbin Hu
Wenbin Hu
Citations: 31
h-index: 2
Weizhen Wang
Weizhen Wang
Citations: 10
h-index: 2

검증 가능한 보상을 사용한 강화 학습(RLVR)은 자동으로 검증 가능한 결과 신호를 사용하여 대규모 언어 모델의 추론 능력을 향상시키지만, 이러한 신호는 일반적으로 희소하고 시퀀스 수준에서만 제공됩니다. 온-폴리시 자체 증류(OPSD)는 학생이 방문하는 접두사에서 권한 있는 교사를 활용하여 토큰 수준의 분포적 감독을 제공함으로써 이러한 희소성을 완화합니다. 이 밀집적인 감독은 신호의 희소성을 줄여주지만, 기존 OPSD는 여전히 롤아웃의 시간 구조를 충분히 활용하지 못한다는 것을 발견했습니다. 기존 OPSD는 각 로컬 발산에 동일한 계수를 할당하는데, 이는 위치나 발생한 발산 시퀀스에 관계없이 적용됩니다. 온-폴리시 자동 회귀 생성에서 동일한 발산 크기는 서로 다른 불일치 이력(teacher와 student 간의 불일치)을 반영할 수 있습니다. 로컬 스칼라 값만으로는 이러한 시간적 맥락을 구별할 수 없으므로, 기존 OPSD는 실현된 불일치 시퀀스에 따라 토큰 수준의 가중치를 조정할 수 없습니다. 이러한 제한 사항을 해결하기 위해, 우리는 발산 적응형 감독 범위(DASH)를 제안합니다. DASH는 각 로컬 증류 신호와 시퀀스 수준 평균 간의 차이를 적응적 전파 게이트로 매핑하고, 이 게이트를 사용하여 역방향 다단계 집계를 제어합니다. 이를 통해 DASH는 생성 과정에서 로컬 발산이 어떻게 진화하는지에 따라 토큰 수준의 감독 가중치를 조정합니다. 세 가지 모델 크기 범위에서 세 가지 수학적 추론 벤치마크에 대한 실험 결과, DASH가 모든 벤치마크와 모든 모델 크기에 대해 기존 OPSD를 능가한다는 것을 보여주었습니다. DASH는 OPSD가 이미 계산하는 교사와 학생의 분포를 재사용하므로, 성능 향상을 위해 추가적인 교사 또는 학생 순방향 연산이 필요하지 않습니다.

Original Abstract

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on-policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token-level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence-Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation. By doing so, DASH adjusts token-level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: https://github.com/DBtxy/DASH-OPSD

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!