OmniVR: 손상된 역사 영화 복원을 위한 비디오-오디오 통합 조건부 생성 모델
OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
역사 영화는 흐릿함, 노이즈, 깜빡임, 히스, 클리핑, 음소거 등 다양한 시각 및 청각적 손상을 동시에 겪습니다. 기존 방법들은 이러한 문제를 개별적으로 해결하려 하지만, 이로 인해 품질 격차가 발생하고 모달 간 일관성이 부족합니다. 본 논문에서는 첫 번째 통합 오디오-비디오 생성 복원 모델인 OmniVR을 제안합니다. 220억 개의 파라미터를 가진 오디오-비디오 생성 기반 모델을 활용하여 OmniVR은 복원을 단일화된 다중 모드 DiT 내에서의 조건부 생성을 통해 수행합니다. 저품질 비디오 및 오디오는 잠재적 조건으로 인코딩되고, 고정된 복원 프롬프트와 결합되어 하나의 조정된 목표 하에서 시각 구조, 시간 동역학 및 음향 디테일을 복원하기 위해 함께 노이즈를 제거합니다. 본 연구에서는 다음과 같은 세 가지 핵심 설계를 통해 이러한 적응을 가능하게 했습니다. (1) 인터넷에서 수집한 데이터로부터 실제 오래된 필름의 특성을 시뮬레이션하는 통합 오디오-비디오 손상 파이프라인, (2) 생성적 사전 지식을 최대한 유지하면서 텍스트-오디오-비디오(T2AV)를 오디오-비디오-오디오-비디오(AV2AV)로 변환하는 아키텍처 보존 프롬프트 어닐링, (3) 장편 비디오 추론 및 음향 충실도를 위한 첫 번째 프레임 이미지-비디오(I2V) 앵커링과 손실 재가중치 부여 및 파형 감독. 또한 OmniVRBench를 제안합니다. 이는 시각 품질, 오디오 품질, 시간적 일관성 및 오디오-시청각 동기화 측면에서 실제 역사 영상 200개를 사용하여 오디오-비디오 복원을 평가하는 최초의 벤치마크입니다. OmniVR은 모든 여섯 가지 시각적 지표에서 기존 방법보다 뛰어난 성능을 보이며, 최상의 오디오 품질을 달성하고 자연스러운 색상화를 구현합니다. 이는 세 가지 측면을 동시에 해결하는 첫 번째 방법입니다. 코드 및 가중치는 공개될 예정입니다. 프로젝트 페이지: https://xin1u.github.io/OminiVR_PAGE/
Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. We present OmniVR, the first joint audio-video generative restoration model. Built upon a 22B-parameter audio-video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low-quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio-video degradation pipeline that simulates real old-film characteristics from Internet-collected data; (2) an architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio-video restoration across visual quality, audio quality, temporal consistency, and audio-visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization---the first method to jointly address all three aspects. Code and weights will be publicly released. Project Page: https://xin1u.github.io/OminiVR_PAGE/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.