2608.04224v1 Aug 04, 2026 cs.CV

OmniVR: 손상된 역사 영화 복원을 위한 비디오-오디오 통합 조건부 생성 모델

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

Zihao Fan
Zihao Fan
Citations: 74
h-index: 5
Xin Lu
Xin Lu
Citations: 231
h-index: 9
Jie Huang
Jie Huang
Citations: 70
h-index: 3
Xueyang Fu
Xueyang Fu
Citations: 832
h-index: 18
Zhengjun Zha
Zhengjun Zha
Citations: 761
h-index: 17
Mingchen Zhong
Mingchen Zhong
Citations: 60
h-index: 3

역사 영화는 흐릿함, 노이즈, 깜빡임, 히스, 클리핑, 음소거 등 다양한 시각 및 청각적 손상을 동시에 겪습니다. 기존 방법들은 이러한 문제를 개별적으로 해결하려 하지만, 이로 인해 품질 격차가 발생하고 모달 간 일관성이 부족합니다. 본 논문에서는 첫 번째 통합 오디오-비디오 생성 복원 모델인 OmniVR을 제안합니다. 220억 개의 파라미터를 가진 오디오-비디오 생성 기반 모델을 활용하여 OmniVR은 복원을 단일화된 다중 모드 DiT 내에서의 조건부 생성을 통해 수행합니다. 저품질 비디오 및 오디오는 잠재적 조건으로 인코딩되고, 고정된 복원 프롬프트와 결합되어 하나의 조정된 목표 하에서 시각 구조, 시간 동역학 및 음향 디테일을 복원하기 위해 함께 노이즈를 제거합니다. 본 연구에서는 다음과 같은 세 가지 핵심 설계를 통해 이러한 적응을 가능하게 했습니다. (1) 인터넷에서 수집한 데이터로부터 실제 오래된 필름의 특성을 시뮬레이션하는 통합 오디오-비디오 손상 파이프라인, (2) 생성적 사전 지식을 최대한 유지하면서 텍스트-오디오-비디오(T2AV)를 오디오-비디오-오디오-비디오(AV2AV)로 변환하는 아키텍처 보존 프롬프트 어닐링, (3) 장편 비디오 추론 및 음향 충실도를 위한 첫 번째 프레임 이미지-비디오(I2V) 앵커링과 손실 재가중치 부여 및 파형 감독. 또한 OmniVRBench를 제안합니다. 이는 시각 품질, 오디오 품질, 시간적 일관성 및 오디오-시청각 동기화 측면에서 실제 역사 영상 200개를 사용하여 오디오-비디오 복원을 평가하는 최초의 벤치마크입니다. OmniVR은 모든 여섯 가지 시각적 지표에서 기존 방법보다 뛰어난 성능을 보이며, 최상의 오디오 품질을 달성하고 자연스러운 색상화를 구현합니다. 이는 세 가지 측면을 동시에 해결하는 첫 번째 방법입니다. 코드 및 가중치는 공개될 예정입니다. 프로젝트 페이지: https://xin1u.github.io/OminiVR_PAGE/

Original Abstract

Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. We present OmniVR, the first joint audio-video generative restoration model. Built upon a 22B-parameter audio-video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low-quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio-video degradation pipeline that simulates real old-film characteristics from Internet-collected data; (2) an architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio-video restoration across visual quality, audio quality, temporal consistency, and audio-visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization---the first method to jointly address all three aspects. Code and weights will be publicly released. Project Page: https://xin1u.github.io/OminiVR_PAGE/

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!