2607.24731v1 Jul 27, 2026 cs.CV

온폴리시 디퓨전 증류에서의 분류기 자유 가이드(Classifier-Free Guidance) 재고

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

Haozhe Wang
Haozhe Wang
Citations: 213
h-index: 5
Fangtai Wu
Fangtai Wu
Citations: 17
h-index: 1
Jiaming Liu
Jiaming Liu
Citations: 822
h-index: 13
Ruihua Huang
Ruihua Huang
Citations: 14
h-index: 1
Jinpeng Yu
Jinpeng Yu
Citations: 196
h-index: 3
Bingnan Li
Bingnan Li
UCSD
Citations: 26
h-index: 3
Haozhong Xiong
Haozhong Xiong
Citations: 0
h-index: 0
Yang Shi
Yang Shi
Citations: 0
h-index: 0

온폴리시 증류(OPD)는 현재 학생 모델이 생성한 경로를 따라 교사 모델을 활용하여 확산 모델을 개선하는 방법이지만, 최신 확산 시스템의 필수 구성 요소인 분류기 자유 가이드(CFG) 환경에서의 동작 방식은 아직 제대로 이해되지 못하고 있습니다. 기존 OPD 방법들은 자연스럽게 속도 정합을 CFG가 적용된 예측에까지 확장하여, 교사와 학생 모델의 가이드된 속도를 직접적으로 일치시키려고 합니다. 본 연구에서는 이러한 목표가 브랜치 수준에서 식별되지 않았음을 보여줍니다. 즉, 양수 및 음수 브랜치의 오류는 가이드된 예측에서 서로 상쇄될 수 있습니다. 두 가지 대조적인 사례를 통해, 공유된 부정 조건 하에서는 기본적인 속도 정합이 효과적임을 확인했습니다. 여기서 양수 및 음수 브랜치 오류가 함께 감소합니다. 그러나 모델의 기본 CFG 스키마가 교사 모델의 음수 브랜치에 학생 모델이 접근할 수 없는 유용한 정보를 포함하고 있는 경우, 이러한 공동 감소 현상은 깨지고, 복합적인 목표는 반대되는 브랜치 오류 동역학을 유발하여 양수 브랜치 오류를 줄이는 동시에 음수 브랜치 오류를 증가시킵니다. 우리는 이 실패 모드를 '음수 브랜치 비대칭(NBA)'이라고 명명했습니다. NBA 문제를 해결하기 위해, 양의 방향 정합(PDM)이라는 새로운 OPD 목표 함수를 제안합니다. PDM은 긍정 예측과 CFG 조건부 방향을 분리하여 제약하는 브랜치 인지적인 목표 함수입니다. 우리는 DPM을 밀집-희소 비디오 제어에 적용했습니다. 여기서 기본적인 가이드된 정합은 추론 가이드 스케일에 매우 민감하지만, 브랜치 인지적인 감독 방법을 통해 더욱 강력하고 효과적인 지식 전달이 가능합니다.

Original Abstract

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!