Poly-OPD: 기능 선택이 가능한 흐름 모델을 위한 이질적인 다중 교사 기반 강화 학습
Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
최첨단 텍스트-이미지 생성 모델들은 종종 상호 보완적인 강점을 가지고 있습니다. 예를 들어, 어떤 모델은 사용자의 선호도에 부합하는 미적 감각에서 뛰어난 반면, 다른 모델은 복합적인 지시사항을 더 정확하게 따르는 경향이 있습니다. 그러나 이러한 모델들의 오토인코더 및 노이즈 스케줄의 차이로 인해 이러한 강점들을 모델 간에 이전하기 어렵습니다. 본 논문에서는 Poly-OPD라는 프레임워크를 제시합니다. 이 프레임워크는 다양한 교사 모델의 상호 보완적인 강점을 단일하고 효율적인 흐름 매칭 기반 학생 모델로 통합합니다. 서로 호환되지 않는 잠재 공간을 연결하기 위해, Poly-OPD는 픽셀 브리지를 통해 온폴리시 증류를 수행합니다. 각 학생 모델이 생성한 이미지는 선택된 교사 모델의 인코더에 의해 다시 인코딩되고, 교사의 노이즈 스케줄과 일치하는 노이즈 레벨로 정제됩니다. 이렇게 얻어진 타겟은 고정된 DINOv2 공간에서 학생 모델과 매칭되어, 서로 호환되지 않는 잠재 공간 간의 학습을 가능하게 합니다. Poly-OPD는 교사 모델 간의 상호 간섭 없이도 다양한 기능을 유지하기 위해 그래디언트 호환성 진단을 사용하여 어댑터를 구성합니다. 어텐션 LoRA 모듈은 모든 교사 모델에서 공유되는 반면, 피드 포워드 어댑터는 각 교사 모델에 특화되어 있습니다. 증류 과정에서, 학생 모델이 아직 교사 모델에 미치지 못하는 복합적인 카테고리에 대해 더 많은 학습을 수행하도록 설계된 가중치를 부여하는 커리큘럼을 사용합니다. 각 격차가 줄어들수록, 학습은 남아 있는 격차가 큰 카테고리로 이동합니다. Poly-OPD는 FLUX.1-dev 및 Z-Image를 2.5B SD3.5-Medium 학생 모델로 증류하여 GenEval 점수를 67.3에서 73.3으로 향상시켰으며, 이는 더 큰 교사 모델보다 우수한 성능입니다. 또한 DrawBench HPSv3 점수를 9.34에서 11.35로 향상시켜, 스위치 가능한 하나의 모델 내에서 다양한 강점을 통합했습니다.
Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to transfer these strengths across models. In this paper, we present Poly-OPD, a framework that can consolidate complementary strengths of heterogeneous teachers into a single compact flow-matching student. To bridge the incompatible latent spaces of different teachers, Poly-OPD performs on-policy distillation through a pixel bridge. Each student-generated image is re-encoded by a selected teacher's encoder and refined from a noise level matched by magnitude under the teacher's noise schedule. The resulting target is further matched to the student in frozen DINOv2 space, enabling supervision across incompatible latent spaces. To retain complementary capabilities without cross-teacher interference, Poly-OPD uses a gradient compatibility diagnostic to organize its adapters: attention LoRA modules are shared across teachers, whereas feed-forward adapters remain teacher-specific. During distillation, a gap-aware curriculum devotes more training to compositional categories where the student still falls short of the teacher. As each gap narrows, training shifts toward categories with larger remaining gaps. By distilling FLUX.1-dev and Z-Image into a 2.5B SD3.5-Medium student, Poly-OPD improves GenEval from 67.3 to 73.3, surpassing both larger teachers, and raises DrawBench HPSv3 from 9.34 to 11.35, consolidating both strengths within a switchable model.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.