OmniDiT: 디퓨전 트랜스포머를 활용한 통합 가상 착용 프레임워크 확장
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
가상 착용(VTON) 및 가상 제거(VTOFF) 기술이 빠르게 발전하고 있지만, 기존 VTON 방법은 미세한 디테일 보존, 복잡한 장면으로의 일반화, 복잡한 파이프라인, 효율적인 추론 등의 어려움에 직면하고 있습니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 디퓨전 트랜스포머 기반의 통합 가상 착용 프레임워크인 OmniDiT를 제안합니다. OmniDiT는 착용 및 제거 작업을 하나의 통합 모델로 결합합니다. 구체적으로, 우리는 먼저 데이터 생성 파이프라인을 구축하여 지속적으로 데이터를 생성하고, 38만 개 이상의 다양한 고품질 의류-모델-착용 이미지 쌍과 상세한 텍스트 프롬프트를 포함하는 대규모 VTON 데이터셋인 Omni-TryOn을 구축합니다. 또한, 토큰 연결 및 적응형 위치 인코딩을 활용하여 여러 참조 조건을 효과적으로 통합합니다. 긴 시퀀스 연산으로 인한 병목 현상을 완화하기 위해, 우리는 디퓨전 모델에 Shifted Window Attention을 처음으로 도입하여 선형 복잡도를 달성합니다. 로컬 윈도우 어텐션으로 인한 성능 저하를 해결하기 위해, 다중 타임스텝 예측과 정렬 손실을 사용하여 생성 품질을 향상시킵니다. 실험 결과, 다양한 복잡한 장면에서 제안하는 방법은 모델 기반 및 모델 프리 VTON, VTOFF 작업 모두에서 최상의 성능을 달성했으며, 모델 기반 VTON 작업에서는 현재 최고 성능 수준과 비교 가능한 성능을 보였습니다.
Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization to complex scenes, complicated pipeline, and efficient inference. To tackle these problems, we propose OmniDiT, an omni Virtual Try-On framework based on the Diffusion Transformer, which combines try-on and try-off tasks into one unified model. Specifically, we first establish a self-evolving data curation pipeline to continuously produce data, and construct a large VTON dataset Omni-TryOn, which contains over 380k diverse and high-quality garment-model-tryon image pairs and detailed text prompts. Then, we employ the token concatenation and design an adaptive position encoding to effectively incorporate multiple reference conditions. To relieve the bottleneck of long sequence computation, we are the first to introduce Shifted Window Attention into the diffusion model, thus achieving a linear complexity. To remedy the performance degradation caused by local window attention, we utilize multiple timestep prediction and an alignment loss to improve generation fidelity. Experiments reveal that, under various complex scenes, our method achieves the best performance in both the model-free VTON and VTOFF tasks and a performance comparable to current SOTA methods in the model-based VTON task.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.