UniVVT: 고품질 비디오 가상 착용을 위한 통합 엔드 투 엔드 프레임워크
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
비디오 가상 착용(VVT)은 대상 의류를 입은 사람의 영상을 합성하면서 개인의 특징, 움직임 및 장면의 역동성을 유지합니다. 기존의 VVT 방식은 주로 마스크 기반의 비디오 채우기 기술을 사용하며, 인간 파싱, 자세 추정 및 의류 변형을 위한 별도의 모듈에 의존합니다. 이러한 다단계 설계는 구현을 복잡하게 만들 뿐만 아니라, 명시적인 기하학적 정보에서의 오류가 생성된 영상으로 지속적으로 전파되는 문제를 야기합니다. 본 논문에서는 VVT를 의미론적으로 조건부 비디오 생성을 문제로 재정의하는 통합 엔드 투 엔드 프레임워크인 UniVVT를 제시합니다. UniVVT는 추론 과정에서 마스크, 자세 및 변형 모듈을 제거합니다. 핵심 구성 요소인 장면-작업 인지 모델은 멀티모달 대규모 언어 모델(Multimodal Large Language Model)을 기반으로 원본 영상, 대상 의류 및 작업 지침을 압축된 작업 관련 잠재 토큰으로 공동 인코딩하여, 어떤 정보를 어디에 어떻게 전송해야 하는지를 암묵적으로 파악합니다. 경량화된 의미론적 브릿지는 이러한 토큰을 확산 기반 비디오 생성기의 조건부 공간과 일치시켜 일관성 있는 의류 전송을 가능하게 합니다. 다양한 구성 요소 간의 견고한 연결을 위해, 의미론적 정렬, 공동 작업 적응 및 유연한 해상도 개선을 포함하는 세 단계의 점진적인 학습 전략을 고안했습니다. 광범위한 실험 결과는 UniVVT가 여러 벤치마크에서 최고 수준의 성능을 달성하며, 암묵적인 의미론적 지침이 엔드 투 엔드 가상 착용을 위한 불안정한 기하학적 전처리 방식에 대한 간단하고 효과적인 대안임을 입증합니다.
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.