CollectionLoRA: 멀티 티처 온폴리시 증류를 통한 단일 LoRA에 50가지 효과 통합
CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation
맞춤형 이미지 편집은 사전 학습된 확산 모델에 제한적인 쌍으로 이루어진 데이터셋을 사용하여 특정 시각적 효과를 부여하는 것을 목표로 하며, 일반적으로 로우-랭크 어댑테이션(LoRA) 기술이 사용됩니다. 원하는 효과의 수가 증가함에 따라, 수많은 효과 LoRA들을 저장하고 동적으로 불러오는 것은 상당한 배포 오버헤드를 발생시킵니다. 또한, 현재의 파이프라인은 이러한 효과 LoRA들을 빠른 생성을 위해 가속화 모듈과 함께 사용하는 경우가 많으며, 이는 심각한 매개변수 간섭을 유발하여 개념 왜곡 및 스타일 저하를 초래합니다. 본 논문에서는 CollectionLoRA라는 멀티 티처 온폴리시 증류 프레임워크를 제안합니다. 이 프레임워크는 최대 50가지의 다양한 효과 LoRA로부터의 정보를 단일 LoRA로 통합하며, 동시에 몇 단계 만에 생성하는 능력을 제공합니다. 이는 근본적으로 특징 간섭 문제를 해결하고 배포 비용을 크게 줄입니다. 구체적으로, 본 방법은 (i) 학습 과정에서 모델이 데이터 소스를 무작위로 전환하도록 하는 확률적 이중 스트림 라우팅 메커니즘을 도입하여, 예측하지 못한 시나리오에서의 일반화 능력을 향상시킵니다; (ii) 프롬프트 공간 내에서 개념 분리를 달성하기 위한 비대칭 직교 프롬프팅 전략을 사용합니다; (iii) 교사 모델과 학생 모델 간의 분포 차이를 줄이기 위해 세분화된 증류 목표 함수를 적용합니다. 광범위한 실험 결과는 CollectionLoRA가 모든 맞춤형 효과와 빠른 생성 기능을 단일 LoRA로 통합하여 배포 오버헤드를 줄이는 동시에, 독립적으로 학습된 교사 모델과 동등하거나 더 나은 수준의 개념 충실도를 달성함을 보여줍니다.
Customized image editing aims to equip pre-trained diffusion models with specific visual effects using limited paired data, typically via Low-Rank Adaptation (LoRA). As the number of desired effects grows, storing and dynamically loading numerous these effect LoRAs significantly increases deployment overhead. Furthermore, current pipelines typically cascade these effect LoRAs with acceleration modules for fast generation, which triggers severe parameter interference and results in concept bleeding and style degradation. We propose CollectionLoRA, a multi-teacher on-policy distillation framework capable of distilling the concepts of up to 50 different effect LoRAs along with few-step generation capabilities into a single LoRA. This fundamentally resolves the feature interference issue and significantly reduces deployment costs. Specifically, the method introduces (i) a Probabilistic Dual-Stream Routing mechanism that enables the model to randomly switch between data sources during training, effectively enhancing its generalization in unseen scenarios; (ii) an Asymmetric Orthogonal Prompting strategy to achieve concept isolation within the prompt space; (iii) a Coarse-to-Fine Distillation Objective to mitigate the distribution gap between the teacher and student models. Extensive evaluations show that CollectionLoRA distills all customized effects and few-step generation into a single LoRA, reducing deployment overhead while achieving concept fidelity comparable to or better than independently trained teacher models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.