SPOT: 희소 탐색 및 결과 보정을 통한 온폴리시 증류
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
온폴리시 증류(OPD)는 학생이 생성한 경로에 대해 밀집된 교사 모델의 지침을 제공하지만, 일반적인 역-KL 훈련 방식은 다른 가능한 연속성에 충분한 확률을 할당하지 못할 수 있습니다. 교사 모델의 엔트로피만으로는 불확실성이 몇 가지 가능성 있는 다음 토큰에 집중되어 있는지, 아니면 긴 확률 분포로 분산되어 있는지, 그리고 학생 모델이 이미 이러한 후보들을 잘 표현하고 있는지 여부를 알 수 없습니다. 또한, 지역적인 교사 모델 확률은 하위 작업에서의 성공을 예측하지 못할 수도 있습니다. 본 논문에서는 두 가지 연관된 결정, 즉 탐색 위치와 증류 대상에 대한 결정을 획득-탐색-활용 절차를 통해 해결하는 희소 탐색 및 결과 보정 온폴리시 증류(SPOT) 방법을 제안합니다. 획득 단계에서는 위치 수준의 점수를 사용하여 정규화된 교사 모델 엔트로피, 소수의 상위 $k$ 후보가 갖는 확률 질량, 그리고 학생-교사 모델 간 불일치를 결합하여 제한된 탐색 예산을 할당합니다. 탐색 단계에서는 SPOT이 검증기(verifier)의 점수를 통해 교사 모델이 제안한 후보들을 평가합니다. 활용 단계에서는 이러한 결과들이 하위 작업에서의 더 나은 성능을 보이는 후보들에게 가중치를 부여하면서도 교사 모델 분포에 고정된, 닫힌 형태의 KL 정규화된 목표를 생성합니다. 다양한 학생 모델 및 추론 벤치마크에서 수행한 광범위한 실험 결과는 SPOT이 추론 성능을 향상시키면서도 솔루션 품질과 범위를 균형 있게 유지하는 데 효과적임을 보여줍니다.
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-$k$ candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.