2606.22830v1 Jun 22, 2026 cs.AI

증거를 찾는 방법: 온-폴리시 추론 증류를 위한 의사 결정 지원 토큰 발견

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation

Xunliang Cai
Xunliang Cai
Citations: 74
h-index: 5
Yueqing Sun
Yueqing Sun
Citations: 32
h-index: 4
Qi Gu
Qi Gu
Citations: 85
h-index: 5
Yuxin Liu
Yuxin Liu
Citations: 9
h-index: 2
Zhuowen Han
Zhuowen Han
Citations: 27
h-index: 3
Zhengxi Lu
Zhengxi Lu
Citations: 290
h-index: 6
Jinwei Xiao
Jinwei Xiao
Citations: 46
h-index: 3
Zhiyuan Yao
Zhiyuan Yao
Citations: 7
h-index: 2
Wentao Chen
Wentao Chen
Citations: 192
h-index: 3

온-폴리시 증류는 밀집된 토큰 레벨의 감독을 통해 추론 능력을 전달하지만, 전달 가능한 신호의 본질은 명확하지 않습니다. 우리는 추론 체인이 두 가지 유형의 지식을 포함하며, 이러한 지식은 서로 다른 발견 메커니즘이 필요하다는 것을 발견했습니다. 첫째, 의사 결정(어디에서 분기해야 하는지)은 학생 모델의 불확실성을 통해 드러나며, 둘째, 증거(의사 결정을 정당화하는 중간 단계)는 학생 모델이 확신하지만 틀린 위치에 숨겨져 있습니다. 현재 방법은 의사 결정만 캡처하며, 증거 토큰에 포함된 중요한 지식은 여전히 전달되지 않습니다. 우리는 DEAR(Decision-Evidence Aware Reasoning Distillation, 의사 결정-증거 인식 추론 증류)를 제안합니다. DEAR은 먼저 학생 모델의 엔트로피를 통해 의사 결정을 식별한 다음, 의사 결정 기준점과의 은닉 상태 코사인 유사성을 통해 해당 의사 결정을 뒷받침하는 증거를 발견합니다. 또한 교사-학생 모델 간의 발산도를 활용하여 지식 격차가 가장 큰 부분을 우선적으로 학습합니다. 수학 및 코드 벤치마크에서 세 가지 학생-교사 구성에 대해 DEAR은 표준 OPD(On-Policy Distillation)보다 일관되게 우수한 성능을 보였으며, 특히 경쟁적인 수학 문제에서는 최대 +2.5pp, 코드 생성에서는 +5.7pp의 향상을 달성했습니다.

Original Abstract

On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chains contain two types of knowledge that require different discovery mechanisms: decisions (where to branch), which surface through student uncertainty, and evidence (intermediate steps that justify decisions), which hides in positions where the student is confident yet wrong. Current methods capture only decisions; the substantive knowledge in evidence tokens remains untransferred. We propose DEAR(Decision-Evidence Aware Reasoning Distillation), which first identifies decisions via student entropy, then discovers their supporting evidence through hidden-state cosine similarity to decision anchors, boosted by teacher-student divergence to prioritize the largest knowledge gaps. Across three student-teacher configurations on math and code benchmarks, DEAR consistently outperforms standard OPD, with up to +2.5pp on competition math and +5.7pp on code generation.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!