2608.04701v1 Aug 05, 2026 cs.CV

UniWorld-View: 비디오 디퓨전 모델을 이용한 광범위 시야 합성

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

Yonghong Tian
Yonghong Tian
Citations: 939
h-index: 9
Li Yuan
Li Yuan
Citations: 733
h-index: 10
Haiyang Zhou
Haiyang Zhou
Citations: 60
h-index: 3
Chaoran Feng
Chaoran Feng
Citations: 382
h-index: 7
Xunyu Zhou
Xunyu Zhou
Citations: 0
h-index: 0
Wangbo Yu
Wangbo Yu
Citations: 738
h-index: 12

소셜 미디어에 존재하는 다양한 단안 동영상 및 이미지는 몰입형 콘텐츠 제작에 귀중한 자료를 제공하며, 이러한 제한적인 데이터를 기반으로 새로운 시점을 생성하면 사용자 경험을 크게 향상시킬 수 있습니다. 그러나 입력 데이터가 매우 제한적인 상황에서 사실적이고 기하학적으로 일관된 시점을 생성하고 정확한 카메라 제어를 수행하는 것은 여전히 어려운 과제입니다. NeRF 및 3D Gaussian Splatting (3DGS)과 같은 재구성 기반 방법은 희소한 입력 데이터에 의해 성능이 크게 저하되며, 명시적인 폐색 처리 기능을 제공하지 못합니다. 생성 모델은 데이터 요구 사항을 완화하지만, 부정확하거나 암묵적인 기하학적 지침으로 인해 광범위 시야 합성에는 어려움이 있습니다. 이러한 한계를 극복하기 위해, 본 논문에서는 단안 입력으로부터 제어 가능한 광범위 새로운 시점 합성을 위한 통합 프레임워크인 UniWorld-View를 소개합니다. UniWorld-View는 명시적인 3D 지침과 생성적 디퓨전 모델링을 결합하여 정확한 카메라 제어를 가능하게 하고, 기하학적으로 일관된 시점 생성을 지원합니다. 기하학적 지침은 폐색을 고려한 포인트 클라우드 렌더링 전략을 통해 얻어지며, 가시성 모호성을 해결하고 디퓨전 기반 합성을 위한 정확한 사전 정보를 제공합니다. 본 연구는 강력한 비디오 디퓨전 모델과 이 렌더링 전략을 결합하여 극단적인 카메라 움직임 및 광범위 시야 변화에서도 고품질의 새로운 시점 생성이 가능하며, 추가적으로 다중 뷰 동영상을 제공하여 후속 동적 3DGS 재구성에 활용될 수 있습니다. WorldScore 벤치마크 및 제로샷 NVS 벤치마크에서의 실험 결과는 UniWorld-View가 제어성, 기하학적 일관성 및 시각적 충실도 측면에서 효과적임을 보여줍니다.

Original Abstract

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!