2608.05776v1 Aug 06, 2026 cs.CV

Vorch-Director: 노이즈 인지 오류 수정 기반의 인터랙티브 세계 스토리 모델

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Xin Ma
Xin Ma
Citations: 807
h-index: 5
Qi Liu
Qi Liu
Citations: 11
h-index: 2
Yang Ding
Yang Ding
Citations: 0
h-index: 0
Siqian Yang
Siqian Yang
Citations: 32
h-index: 4
Gang Yue
Gang Yue
Citations: 412
h-index: 3
Lin Ma
Lin Ma
Citations: 56
h-index: 2
Yidi Wu
Yidi Wu
Citations: 32
h-index: 4
Lisai Zhang
Lisai Zhang
Citations: 9
h-index: 2
Yaohui Wang
Yaohui Wang
Citations: 18
h-index: 2
Jingyuan Chen
Jingyuan Chen
Citations: 2,126
h-index: 14

자기 회귀(autoregressive) 방식은 짧은 구간 생성기를 반복적으로 확장하여 미세 규모의 오디오-비디오 생성을 위한 자연스러운 방법을 제공합니다. 하지만, 모델은 깨끗한 원본 데이터로 훈련되지만, 추론 과정에서는 모델 스스로 생성한 데이터를 기반으로 하기 때문에 누적된 오류로 인해 객체 일관성(identity drift)이 발생하고, 과도한 평활화(over-smoothing)가 나타나며 오디오-비디오 동기화 문제가 발생합니다. 최근 연구들은 예측 잔차(prediction residuals)를 인공적인 노이즈로 활용하여 이러한 불일치를 줄이려고 시도했지만, 우리는 잔차 수정의 효과가 잔차가 생성되는 노이즈 수준에 크게 의존한다는 것을 확인했습니다. 본 논문에서는 각 잔차와 그 잔차가 발생한 노이즈 수준을 연결하는 노이즈 수준 인지 잔차 수정 전략인 Vorch-Director를 제안합니다. Vorch-Director는 훈련 과정에서 일치하는 노이즈 레벨의 잔차를 주입하여, 주입된 오류를 디노이징(denoising) 프로세스와 연계함으로써 더욱 현실적인 자기 회귀 시퀀스를 생성하며 효율적인 teacher-forcing 훈련 방식을 유지합니다. 오디오-비디오 LTX-2 diffusion transformer를 기반으로 구축된 Vorch-Director는 또한, 과거 비디오, 참조 이미지 및 대상 비디오를 구별하기 위한 태스크 임베딩을 도입하여 장기 생성을 위한 통합적 조건 부여를 가능하게 합니다. 깨끗한 조건부 입력(conditioning sink)과 혼합 작업 훈련(mixed-task training)과 함께 Vorch-Director는 멀티샷, 멀티 서브젝트, 참조 기반 오디오-비디오 장편 비디오 생성 기능을 지원합니다. 우리는 ST-Bench 데이터셋에서 Vorch-Director를 평가하고, 품질 변화 및 장거리 일관성을 측정하기 위한 새로운 오디오-비디오 장기 생성 벤치마크를 소개합니다. 광범위한 실험 결과는 Vorch-Director가 기존 모델보다 향상된 안정성과 오디오-비디오 충실도를 제공함을 보여줍니다.

Original Abstract

Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated video and audio. However, models are trained on clean ground-truth histories, while inference relies on their own generated histories, where accumulated errors cause identity drift, over-smoothing, and audio-visual desynchronization. Recent methods reduce this mismatch by reusing prediction residuals as synthetic corruption, but we observe that the effectiveness of residual correction critically depends on the flow-matching noise level at which residuals are produced. We propose Vorch-Director, a noise-level-aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training. By aligning injected errors with the denoising process, Vorch-Director produces more realistic autoregressive histories while retaining efficient teacher-forcing training. Built on the audio-visual LTX-2 diffusion transformer, Vorch-Director further introduces task embeddings to distinguish historical video, reference images, and target video, enabling unified conditioning for long-horizon generation. Together with a clean conditioning sink and mixed-task training, Vorch-Director supports multi-shot, multi-subject, reference-guided audio-visual long-video generation. We evaluate Vorch-Director on ST-Bench and introduce a new long-horizon audio-visual benchmark with metrics for quality drift and long-range consistency. Extensive experiments demonstrate improved stability and audio-visual fidelity over strong baselines.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!