Vorch-Omni: 시각 및 청각의 다중 작업 통합
Vorch-Omni: Multi-Task Orchestration of Sight and Sound
최근 생성 비디오 모델링 기술 발전은 다양한 콘텐츠 생성, 참조 기반 합성, 확장 및 편집 기능을 가능하게 했지만, 기존 접근 방식은 종종 단편적인 작업별 모델에 의존합니다. 일반적인 모델은 생성할 내용, 보존할 요소 또는 지침으로 사용할 요소를 결정하기 위해 이질적인 대상, 소스 및 참조 신호를 구별해야 하며, 동시에 작업 간의 간섭을 줄여야 합니다. 오디오-비주얼 생성은 다양한 조건 설정과 모달리티 간 출력 구성이 추가되어 이러한 과제를 더욱 심화시킵니다. 본 논문에서는 임의 조건에서 임의 출력을 생성하는 통합 다중 작업 프레임워크인 Vorch-Omni를 제안합니다. Vorch-Omni는 비디오 및 오디오 신호를 조건 입력 또는 생성 대상으로 유연하게 처리할 수 있습니다. 토큰 수준의 조건 마스크와 작업 식별자는 대상, 소스 콘텐츠 및 참조를 구별하고, 위치 유형은 시간적 맥락과 독립적인 조건을 분리합니다. Vorch-Omni는 의미론적 및 구조적 정보를 캡처하기 위해 상호 보완적인 시각 조건 설정 경로를 사용합니다. 비전-언어 모델은 샘플링된 프레임을 텍스트 지침과 함께 해석하고, 비디오 VAE는 조건을 잠재 토큰으로 인코딩하여 직접적인 가이드 역할을 수행합니다. 또한, 다양한 시간적으로 정렬된 오디오-비주얼 클립을 선별하고 구조화된 캡션 및 메타데이터를 생성하며, 이질적인 작업 분포를 균형 있게 구성하기 위한 분산 데이터 파이프라인을 구축했습니다. Vorch-Omni는 작업별 아키텍처 변경 없이 단일 플로우 매칭 디퓨전 트랜스포머를 기반으로 하며, 텍스트-비디오 변환, 텍스트-오디오-비디오 변환, 이미지 및 참조 조건부 생성, 시간 확장, 오디오 기반 생성, 비디오 변환, 오디오-비주얼 편집 등 10개 이상의 작업을 지원합니다. 이러한 통합 프레임워크는 범용 오디오-비주얼 생성 및 조작을 위한 확장 가능한 기반을 제공합니다.
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.