Vorch-IR: 장편 통합 다중 모드 동영상 내 인물 교체 비디오 생성
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
동영상 인물 교체는 주 동영상의 움직임, 표정 및 시간 구조를 유지하면서 하나 이상의 대상의 정체를 이전하는 것을 목표로 합니다. 기존 방법은 주로 단일 인물 환경을 대상으로 하며, 종종 마스크나 자세 표현과 같은 작업별 구조 제어를 필요로 하여 일반적인 다중 모드 편집 시스템에서의 유연성을 제한합니다. 다중 인물 교체에 대한 발전은 쌍으로 이루어진 학습 데이터의 부족으로 더욱 제한됩니다. 본 논문에서는 단일 및 이중 인물 교체를 지원하며, 선택적으로 배경 교체가 가능한 통합 프레임워크인 Vorch-IR을 제시합니다. LTX2를 기반으로 구축된 Vorch-IR은 주 동영상, 색인화된 참조 이미지 및 텍스트 편집 지시 사항에 대해 동시에 조건을 부여합니다. 참조 이미지는 주 동영상의 자세, 레이아웃 또는 공간 구성과 일치할 필요가 없으며, 대상 또는 배경 참조로서의 역할은 지시 사항을 통해 지정됩니다. 밀집된 시각적 조건은 자체 주의 메커니즘을 통해 융합되고, 시각-언어 맥락은 크로스 어텐션을 통해 의미론적 대응을 확립합니다. 또한, 모든 네 가지 편집 설정에 대한 쌍으로 이루어진 지도 학습 데이터를 자동으로 생성하는 데이터 구축 파이프라인을 개발했습니다. 자동 측정 및 인간 평가를 사용한 실험 결과는 다양한 시나리오에서 강력한 정체성 유지, 동작 충실도 및 시간적 일관성을 보여줍니다. 또한, 시간 중첩 추론 전략을 통해 짧은 클립 모델을 사용하여 반복적인 연결 없이 최대 분 길이의 비디오 생성이 가능합니다.
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single-person settings and often require task-specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi-person replacement is further constrained by the scarcity of paired training data. We present Vorch-IR, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch-IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self-attention, while a vision-language context establishes semantic correspondence through cross-attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short-clip model to minute-long generation without autoregressive continuation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.