2608.01735v1 Aug 03, 2026 cs.AI

DAPD: 이중 고정 정책 증류

DAPD: Dual-Anchored Policy Distillation

Encheng Su
Encheng Su
Citations: 60
h-index: 4
Jianyu Wu
Jianyu Wu
Citations: 57
h-index: 5
Chen Tang
Chen Tang
Citations: 46
h-index: 4
Yizhou Wang
Yizhou Wang
Citations: 55
h-index: 3
Shixiang Tang
Shixiang Tang
Citations: 47
h-index: 3

온라인(자기) 증류(OPSD)는 언어 모델의 추가 학습에 점점 더 많이 사용되고 있습니다. OPSD는 교사 모델에게 특권 정보를 제공하여 성능을 향상시키지만, '특권 환상'을 유발할 수 있습니다. 이는 학생 모델이 추론 시점에서 사용할 수 없는 특권 정보에 의존적인 동작을 학습하고, 마치 훈련 시점의 특권 정보가 여전히 사용 가능한 것처럼 행동하게 되어 결국 성능 저하를 초래합니다. 본 논문에서는 OPSD의 실패 원인을 추론 시점에서 교사 모델과 학생 모델 간의 정보 불균형으로 규명했습니다. 이러한 불균형을 해결하기 위해, 우리는 이중 고정 정책 증류(DAPD)라는 통합 프레임워크를 제안합니다. DAPD는 두 가지 수준의 고정을 제공합니다. 먼저, 이중 경로 고정(DPA)은 자기 조건화된 브릿지를 도입하고, 참조 행동과 롤아웃 행동을 두 개의 일치하는 정보 경로에 맞춰 정렬하여, 특권 정보에 의존적인 동작이 추론 시점 학생 모델로 전달되는 것을 방지합니다. 또한, 이중 소스 고정(DSA)은 이러한 경로를 참조에서 롤아웃 방향과 롤아웃에서 참조 방향 모두에 적용하여, 특권 참조 가이드에 대한 의존성을 줄이면서도 정확성 감독을 유지합니다. 광범위한 실험 결과는 DAPD가 특권 환상을 크게 완화하며, Qwen3-4B 모델에서 평균적으로 +2.00점의 성능 향상을 보인다는 것을 보여줍니다. 특히, 이러한 성능 향상은 모델 크기에 관계없이 지속되며, 4B 모델에서는 +2.69점, 32B 모델에서는 +2.78점의 개선 효과를 나타냅니다.

Original Abstract

On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege illusion: the student learns privilege-dependent behavior it cannot reproduce from its inference-time context, yet behaves as if the training-time privileged information remained available, ultimately degrading performance. In this paper, we identify information asymmetry between the privileged teacher and the student at inference as the root cause of this failure in OPSD. To resolve this asymmetry, we propose Dual-Anchored Policy Distillation (DAPD), a unified framework with two levels of anchoring. Dual-Path Anchoring (DPA) introduces a self-conditioned bridge and aligns reference and rollout behavior along two matched-information paths, preventing privilege-dependent behavior from being transferred to the inference-time student. Dual-Source Anchoring (DSA) applies these paths in both reference-to-rollout and rollout-to-reference directions, reducing reliance on privileged reference guidance while preserving correctness supervision. Extensive experiments show that DAPD significantly alleviates privilege illusion, outperforming OPSD on Qwen3-4B by +2.00 points on average across tasks. Notably, its gains persist across scales, reaching +2.69 at 4B and +2.78 at 32B.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!