2608.03632v1 Aug 04, 2026 cs.AI

교사가 잘못 인도할 때: 오해를 줄이는 온라인 정책 증류

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Huajun Chen
Huajun Chen
Citations: 616
h-index: 12
Zhuang Xiang
Zhuang Xiang
Citations: 823
h-index: 12
Yongjie Ye
Yongjie Ye
Citations: 355
h-index: 7
Zhou Tao
Zhou Tao
Citations: 5
h-index: 2
Tiankai Li
Tiankai Li
Citations: 0
h-index: 0

온라인 정책 증류(OPD)는 교사의 능력을 학생이 생성한 경로에 대한 밀집된 토큰 수준의 교사 신호를 통해 전달합니다. 최근의 선택적 OPD 방법은 신뢰할 수 있거나, 정보적이거나, 학습 가능성이 높은 신호에 우선순위를 부여하여 이 프로세스를 개선합니다. 그러나 이러한 접근 방식은 언어 모델의 근본적인 문제점을 간과합니다. 즉, 토큰 수준의 판단이 작업별 증거보다는 입력에 독립적인 언어적 선입견, 서식 규칙 또는 고정관념화된 추론 템플릿에 의해 좌우될 수 있다는 점입니다. 우리는 OPD에서 이러한 최적화와 관련된 반면 약하게 입력 기반의 감독을 '오해의 소지가 있는 신호'라고 부르며, 이는 큰 기울기를 발생시키면서도 작업 개선에 거의 기여하지 않을 수 있습니다. 이 문제를 완화하기 위해, 입력 기반으로 오해의 소지가 있는 신호를 식별하고 필터링하는 'Spurious-Signal-Aware On-Policy Distillation(SA-OPD)' 프레임워크를 제안합니다. SA-OPD는 토큰 수준의 증류 신호가 실제로 입력에 의존하는지 여부를 추정하는 경량의 입력 기반 검증 프록시를 도입합니다. 그런 다음, 낮은 입력 기반성과 극단적인 증류 발산을 동시에 나타내는 토큰만 필터링하여 고영향의 오해의 소지가 있는 업데이트를 제거하고 미세하게 조정된 OPD 최적화를 달성합니다. 대규모 언어 모델(LLM) 및 비전-언어 모델(VLM) 환경에서의 광범위한 실험 결과, SA-OPD가 기존의 OPD 방법과 경쟁적인 선택적 방법보다 일관되게 우수한 성능을 보인다는 것을 보여줍니다. 이러한 결과는 입력 기반성을 OPD 감독 신호 선택의 핵심 요소로 확립하고, 오해의 소지가 있는 업데이트를 완화하는 간단하고 효과적인 전략을 제시합니다.

Original Abstract

On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!