2608.05803v1 Aug 06, 2026 cs.CV

Vorch-Omni: 시각 및 청각의 다중 작업 통합

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Peng Shi
Peng Shi
Citations: 51
h-index: 4
Xin Ma
Xin Ma
Citations: 807
h-index: 5
Qi Liu
Qi Liu
Citations: 11
h-index: 2
Vorch Team
Vorch Team
Citations: 0
h-index: 0
Yang Ding
Yang Ding
Citations: 0
h-index: 0
Cong Han
Cong Han
Citations: 19
h-index: 2
Menglin Han
Menglin Han
Citations: 11
h-index: 2
Yuxin Hong
Yuxin Hong
Citations: 27
h-index: 2
Jiebo Hou
Jiebo Hou
Citations: 0
h-index: 0
Zequn Jie
Zequn Jie
Citations: 44
h-index: 3
Siyuan Luo
Siyuan Luo
Citations: 24
h-index: 3
Yinlong Qian
Yinlong Qian
Citations: 94
h-index: 5
Fang Wan
Fang Wan
Citations: 449
h-index: 6
Siqian Yang
Siqian Yang
Citations: 32
h-index: 4
Mingyu Yin
Mingyu Yin
Citations: 20
h-index: 3
Haoran Yu
Haoran Yu
Citations: 0
h-index: 0
Gang Yue
Gang Yue
Citations: 412
h-index: 3
Xiaoyu Chen
Xiaoyu Chen
Citations: 0
h-index: 0
Xiang Li
Xiang Li
Citations: 0
h-index: 0
Jing Liu
Jing Liu
Citations: 24
h-index: 3
Yulei Lu
Yulei Lu
Citations: 94
h-index: 3
Lin Ma
Lin Ma
Citations: 56
h-index: 2
Siqi Wang
Siqi Wang
Citations: 24
h-index: 3
Yidi Wu
Yidi Wu
Citations: 32
h-index: 4
Lisai Zhang
Lisai Zhang
Citations: 9
h-index: 2

최근 생성 비디오 모델링 기술 발전은 다양한 콘텐츠 생성, 참조 기반 합성, 확장 및 편집 기능을 가능하게 했지만, 기존 접근 방식은 종종 단편적인 작업별 모델에 의존합니다. 일반적인 모델은 생성할 내용, 보존할 요소 또는 지침으로 사용할 요소를 결정하기 위해 이질적인 대상, 소스 및 참조 신호를 구별해야 하며, 동시에 작업 간의 간섭을 줄여야 합니다. 오디오-비주얼 생성은 다양한 조건 설정과 모달리티 간 출력 구성이 추가되어 이러한 과제를 더욱 심화시킵니다. 본 논문에서는 임의 조건에서 임의 출력을 생성하는 통합 다중 작업 프레임워크인 Vorch-Omni를 제안합니다. Vorch-Omni는 비디오 및 오디오 신호를 조건 입력 또는 생성 대상으로 유연하게 처리할 수 있습니다. 토큰 수준의 조건 마스크와 작업 식별자는 대상, 소스 콘텐츠 및 참조를 구별하고, 위치 유형은 시간적 맥락과 독립적인 조건을 분리합니다. Vorch-Omni는 의미론적 및 구조적 정보를 캡처하기 위해 상호 보완적인 시각 조건 설정 경로를 사용합니다. 비전-언어 모델은 샘플링된 프레임을 텍스트 지침과 함께 해석하고, 비디오 VAE는 조건을 잠재 토큰으로 인코딩하여 직접적인 가이드 역할을 수행합니다. 또한, 다양한 시간적으로 정렬된 오디오-비주얼 클립을 선별하고 구조화된 캡션 및 메타데이터를 생성하며, 이질적인 작업 분포를 균형 있게 구성하기 위한 분산 데이터 파이프라인을 구축했습니다. Vorch-Omni는 작업별 아키텍처 변경 없이 단일 플로우 매칭 디퓨전 트랜스포머를 기반으로 하며, 텍스트-비디오 변환, 텍스트-오디오-비디오 변환, 이미지 및 참조 조건부 생성, 시간 확장, 오디오 기반 생성, 비디오 변환, 오디오-비주얼 편집 등 10개 이상의 작업을 지원합니다. 이러한 통합 프레임워크는 범용 오디오-비주얼 생성 및 조작을 위한 확장 가능한 기반을 제공합니다.

Original Abstract

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!