STEP-OPD: 확산 모델의 온-폴리시 증류에서 출력 목표와 내부 역학 재고
STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
온-폴리시 증류(OPD)는 여러 개의 특정 작업에 특화된 이미지 생성 모델을 하나의 학생 모델로 통합하는 효과적인 방법으로 자리 잡았습니다. 그러나 기존의 OPD 방법은 주로 학생 모델이 교사 모델의 출력 속도를 맞추도록 최적화하는데, 이는 교사 모델을 최적화 목표의 상한선으로 만듭니다. 출력 수준의 감독만으로는 학생 모델의 블록 단위 표현 변화가 충분히 제약되지 않아, 계층적으로 점진적으로 발전해야 하는 능력을 이전하는 데 어려움이 있습니다. 본 논문에서는 이미지 생성에 대한 온-폴리시 증류 프레임워크인 STEP-OPD를 제안합니다. STEP-OPD는 학생 모델의 학습 목표를 교사 모델을 넘어 확장하고, 내부 표현 변화에 대한 명시적인 제약을 도입합니다. 우리는 교사 모델을 최종 목표로 간주하는 대신, 각 작업별 교사 모델과 공유된 기본 모델 사이의 속도 차이를 추가 학습 방향으로 활용하고, 이 차이의 스케일 조정된 버전을 교사 모델의 속도에 더합니다. 또한, 학생 모델과 교사 모델 간의 표현 변화 방향 및 크기를 일치시켜, 학생 모델이 네트워크 블록 전체에서 표현이 어떻게 점진적으로 변환되는지 학습하도록 합니다. 합성 정렬, 텍스트 렌더링 및 인간 선호도 실험 결과, 제안하는 방법은 기존의 표준 OPD 방법을 꾸준히 개선합니다. 특히, DiffusionOPD의 GenEval 점수를 0.927에서 0.961로 향상시키고, OCR 성능과 모든 선호도 기반 지표를 개선했습니다. 결과적으로 통합된 학생 모델은 세 가지 능력 그룹 모두에서 해당 단일 작업 교사 모델을 능가하며, 이는 출력 외삽이 교사 모델의 한계를 넘어 학습할 수 있도록 한다는 것을 보여줍니다. 또한, 표현 변화 일치는 학생 모델의 내부 변환에 대한 상호 보완적인 지침을 제공합니다.
On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD methods optimize the student mainly to match the teacher's output velocity, making the teacher the upper limit of the optimization objective. While output-level supervision alone leaves the student's blockwise representation evolution underconstrained, which weakens the transfer of capabilities that must be progressively developed across layers. We propose STEP-OPD, an on-policy distillation framework for image generation that extends the student's learning target beyond the teacher and introduces explicit constraints on its internal representation evolution. Instead of treating the teacher as the final target, we use the velocity difference between each task-specific teacher and the shared base model as a direction for further learning and add a scaled version of this difference to the teacher velocity. In addition, we align the direction and magnitude of representation changes between the student and teacher, enabling the student to learn how representations are progressively transformed across network blocks. Experiments on compositional alignment, text rendering, and human preference show that our method consistently improves Standard OPD methods. In particular, it increases the GenEval score of DiffusionOPD from 0.927 to 0.961, while also improving OCR and all preference-based metrics. The resulting unified student surpasses the corresponding single-task teachers across all three capability groups, showing that output extrapolation enables beyond-teacher learning. And representation change alignment provides complementary guidance for the student's internal transformations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.