2606.30626v1 Jun 29, 2026 cs.AI

DOPD: 이중 온라인 정책 증류

DOPD: Dual On-policy Distillation

Kaituo Feng
Kaituo Feng
Citations: 830
h-index: 11
Xiaobin Hu
Xiaobin Hu
Citations: 214
h-index: 4
Xinlei Yu
Xinlei Yu
Citations: 353
h-index: 8
Xiangyu Yue
Xiangyu Yue
Citations: 650
h-index: 9
Shuicheng Yan
Shuicheng Yan
Citations: 584
h-index: 7
Guibin Zhang
Guibin Zhang
Citations: 295
h-index: 6
Qunzhong Wang
Qunzhong Wang
Citations: 50
h-index: 3
Shuai Dong
Shuai Dong
Citations: 36
h-index: 3
Gen Li
Gen Li
Citations: 0
h-index: 0
Qingyi Si
Qingyi Si
Citations: 48
h-index: 3
Kaiwen Tuo
Kaiwen Tuo
Citations: 28
h-index: 2
Yuqi Xu
Yuqi Xu
Citations: 48
h-index: 4
Congcong Wang
Congcong Wang
Citations: 0
h-index: 0
Xiangyu Zeng
Xiangyu Zeng
Citations: 0
h-index: 0
Yang Shi
Yang Shi
Citations: 0
h-index: 0
Jiaqi Wang
Jiaqi Wang
Citations: 484
h-index: 4

온라인 증류(OPD)는 학생 모델이 생성한 경로에 대한 상세한 토큰 수준 신호를 활용하여 우수한 성능 향상을 제공합니다. 증류의 성능을 향상시키기 위해, 가이드 정보(privileged information)를 교수 모델 또는 학생 모델 자체에 주입하는 것은 직관적인 방법입니다. 그러나 이러한 추가 입력은 '특권 환상(privilege illusion)'이라는 잠재적인 문제점을 야기할 수 있습니다. 이는 학생들이 학습해야 할 능력 격차와 모방될 수 있지만 재현될 수 없는 정보 비대칭 격차를 혼동시키는 현상을 의미합니다. 특히, 토큰 수준의 감독 신호는 고유한 불균일성을 가지며, 중요한 능력을 담고 있는 토큰은 극히 일부에 불과합니다. 이러한 문제를 해결하기 위해, 우리는 DOPD라는 장점(advantage)을 고려한 이중 증류 프레임워크를 제안합니다. DOPD는 교수 모델과 학생 모델의 장점 격차와 상대적인 확률을 기반으로, 토큰 수준의 감독 신호를 동적으로 교환합니다. 각 토큰은 교수 또는 학생 모델로부터 서로 다른 강도, 목적, 전략에 따른 감독을 받으며, 이는 실제 능력을 전달하는 동시에 보조 신호를 제공하여 특권 환상을 완화합니다. 대규모 언어 모델(LLM) 및 시각-언어 모델(VLM) 환경에서의 광범위한 실험 결과는 DOPD가 일반적인 OPD 및 다른 방법보다 일관되게 우수한 성능을 발휘함을 보여줍니다. 또한, 안정성, 강건성, 지속 학습 및 이상 데이터셋에 대한 추가적인 결과는 DOPD의 우수성을 입증합니다.

Original Abstract

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities. Each token receives supervision of different strength, objective, and strategy from either teacher or student itself, which transfers credible capability while simultaneously receiving auxiliary signals, to alleviate privilege illusion. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that DOPD consistently outperforms Vanilla OPD and other counterparts. Further results on stability, robustness, continual learning, and out-of-distribution tasks validate its superiority.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!