2608.05745v1 Aug 06, 2026 cs.CV

UniVVT: 고품질 비디오 가상 착용을 위한 통합 엔드 투 엔드 프레임워크

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

Yiheng Zhu
Yiheng Zhu
Citations: 21
h-index: 3
Yushe Cao
Yushe Cao
Citations: 31
h-index: 1
Shikun Feng
Shikun Feng
Citations: 5
h-index: 2
Fei Shen
Fei Shen
Citations: 73
h-index: 2
Haikuo Peng
Haikuo Peng
Citations: 0
h-index: 0
Jianqiang Xia
Jianqiang Xia
Citations: 29
h-index: 2
Dian-xi Shi
Dian-xi Shi
Citations: 5
h-index: 1
Chun Yu
Chun Yu
Citations: 358
h-index: 11

비디오 가상 착용(VVT)은 대상 의류를 입은 사람의 영상을 합성하면서 개인의 특징, 움직임 및 장면의 역동성을 유지합니다. 기존의 VVT 방식은 주로 마스크 기반의 비디오 채우기 기술을 사용하며, 인간 파싱, 자세 추정 및 의류 변형을 위한 별도의 모듈에 의존합니다. 이러한 다단계 설계는 구현을 복잡하게 만들 뿐만 아니라, 명시적인 기하학적 정보에서의 오류가 생성된 영상으로 지속적으로 전파되는 문제를 야기합니다. 본 논문에서는 VVT를 의미론적으로 조건부 비디오 생성을 문제로 재정의하는 통합 엔드 투 엔드 프레임워크인 UniVVT를 제시합니다. UniVVT는 추론 과정에서 마스크, 자세 및 변형 모듈을 제거합니다. 핵심 구성 요소인 장면-작업 인지 모델은 멀티모달 대규모 언어 모델(Multimodal Large Language Model)을 기반으로 원본 영상, 대상 의류 및 작업 지침을 압축된 작업 관련 잠재 토큰으로 공동 인코딩하여, 어떤 정보를 어디에 어떻게 전송해야 하는지를 암묵적으로 파악합니다. 경량화된 의미론적 브릿지는 이러한 토큰을 확산 기반 비디오 생성기의 조건부 공간과 일치시켜 일관성 있는 의류 전송을 가능하게 합니다. 다양한 구성 요소 간의 견고한 연결을 위해, 의미론적 정렬, 공동 작업 적응 및 유연한 해상도 개선을 포함하는 세 단계의 점진적인 학습 전략을 고안했습니다. 광범위한 실험 결과는 UniVVT가 여러 벤치마크에서 최고 수준의 성능을 달성하며, 암묵적인 의미론적 지침이 엔드 투 엔드 가상 착용을 위한 불안정한 기하학적 전처리 방식에 대한 간단하고 효과적인 대안임을 입증합니다.

Original Abstract

Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!