2608.09263v1 Aug 10, 2026 cs.AI

우월한 확률이 자동으로 가치를 의미하지는 않는다: 온-정책 자체 증류에서의 토큰 기여도에 대한 세 가지 검증

Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xuan-Phi Nguyen
Xuan-Phi Nguyen
Citations: 877
h-index: 15
Shafiq Joty
Shafiq Joty
Citations: 816
h-index: 18
Shrey Pandit
Shrey Pandit
University of Texas at Austin
Citations: 254
h-index: 9
Anurag Koul
Anurag Koul
Citations: 772
h-index: 5

결과 검증 도구는 완료된 추론 과정을 평가하지만, 중간 단계의 토큰에는 가중치를 부여하지 않습니다. 특권적 자체 증류(privileged self-distillation)는 훈련 데이터만 사용하여 모델의 자체 출력을 재평가함으로써 이러한 격차를 해소하려고 시도합니다. 그러나 토큰 확률 변화가 자동으로 결과에 대한 기여도를 나타내는 것은 아닙니다. 우리는 다음 세 가지 질문으로 문제를 분리하여 분석합니다: (1) 점수가 더 나은 행동을 잘 반영하는지, (2) 피드백 구성 방식이 비교 대상에 어떤 영향을 미치는지, 그리고 (3) 훈련 손실이 어떤 동작을 강화하는지를 확인해야 합니다. 우리는 이러한 구분을 형식적으로 제시합니다. 특정 추론 과정의 결과에 대한 피드백을 사용하여 점수를 매기는 경우, 해당 내용이 토큰과 평가 맥락 모두를 결정하므로 직접적인 자기 의존성이 발생합니다. 동일 문제에 대한 다른 추론 과정에서 얻은 피드백을 사용하면 이러한 의존성은 제거되지만, 유용한 점수를 보장하지는 않습니다. AIME 2025 데이터셋에 대해 20B 모델로 수행한 실험에서 구현된 추가적인 점수는 무작위 수준(AUC=0.505)과 거의 비슷하며, 길이 조정 후 약간 부정확한 추론 과정을 선호하는 경향이 있었습니다. 쌍을 이룬 비교에서는 결과만을 고려한 제어 그룹이 64.2%의 성능을 보인 반면, 다섯 가지 토큰-점수 변형은 24.2%에서 33.9%의 성능을 보였습니다. 이러한 결과는 확률 신호가 실제로 가치를 나타내는지를 판단하기 전에 점수의 의미, 피드백 구성 방식, 그리고 훈련 동작을 개별적으로 검증해야 함을 시사합니다.

Original Abstract

Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts to fill this gap by rescoring a model's own rollout with training-only information. A token likelihood change, however, is not automatically outcome credit. We separate three questions: whether the score tracks better actions, whether feedback construction changes what is compared, and what behavior the training loss reinforces. We establish these distinctions formally. When a rollout is scored using hindsight feedback written about that same rollout, its content determines both the tokens and the scoring context, creating direct self-dependence. Using feedback from another rollout of the same problem removes this dependence but does not guarantee a useful score. In matched experiments with a 20B model on AIME 2025, the implemented additive score is near chance (AUC=0.505) and slightly favors incorrect traces after length adjustment. In the paired comparison, the outcome-only control records 64.2\%, versus 24.2\%--33.9\% for five token-score variants. The results motivate validating score meaning, feedback construction, and training behavior separately before calling a likelihood signal credit.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!