2604.02828v1 Apr 03, 2026 cs.CV

NavCrafter: 단일 이미지로부터 3D 장면 탐색

NavCrafter: Exploring 3D Scenes from a Single Image

Fangming Liu
Fangming Liu
Citations: 14
h-index: 3
Xueqian Wang
Xueqian Wang
Citations: 21
h-index: 2
H. Duan
H. Duan
Citations: 0
h-index: 0
Peiyu Zhuang
Peiyu Zhuang
Citations: 191
h-index: 3
Yi Liu
Yi Liu
Citations: 5
h-index: 2
Zhengyang Zhang
Zhengyang Zhang
Citations: 1
h-index: 1
Yuxin Zhang
Yuxin Zhang
Citations: 2
h-index: 1
Pengting Luo
Pengting Luo
Citations: 0
h-index: 0

3D 데이터를 직접 획득하기 어렵거나 비용이 많이 드는 경우, 단일 이미지로부터 유연한 3D 장면을 생성하는 것은 매우 중요합니다. 본 논문에서는 NavCrafter라는 새로운 프레임워크를 소개합니다. NavCrafter는 카메라 제어 가능성과 시간-공간 일관성을 갖춘 새로운 시점 비디오 시퀀스를 합성하여 단일 이미지로부터 3D 장면을 탐색합니다. NavCrafter는 풍부한 3D 사전 지식을 캡처하기 위해 비디오 확산 모델을 활용하며, 장면 커버리지를 점진적으로 확장하기 위한 기하학적 인식을 갖춘 확장 전략을 채택합니다. 제어 가능한 다중 시점 합성을 가능하게 하기 위해, 우리는 듀얼 브랜치 카메라 주입 및 어텐션 조절을 통해 다양한 경로를 확산 모델에 적용하는 다단계 카메라 제어 메커니즘을 도입했습니다. 또한, 충돌을 고려한 카메라 경로 계획 알고리즘과 깊이 정렬 기반의 지도, 구조적 정규화 및 개선된 3D Gaussian Splatting (3DGS) 파이프라인을 제안합니다. 광범위한 실험 결과는 NavCrafter가 큰 시점 변화에서도 최첨단 수준의 새로운 시점 합성을 달성하며, 3D 재구성 정확도를 크게 향상시킨다는 것을 보여줍니다.

Original Abstract

Creating flexible 3D scenes from a single image is vital when direct 3D data acquisition is costly or impractical. We introduce NavCrafter, a novel framework that explores 3D scenes from a single image by synthesizing novel-view video sequences with camera controllability and temporal-spatial consistency. NavCrafter leverages video diffusion models to capture rich 3D priors and adopts a geometry-aware expansion strategy to progressively extend scene coverage. To enable controllable multi-view synthesis, we introduce a multi-stage camera control mechanism that conditions diffusion models with diverse trajectories via dual-branch camera injection and attention modulation. We further propose a collision-aware camera trajectory planner and an enhanced 3D Gaussian Splatting (3DGS) pipeline with depth-aligned supervision, structural regularization and refinement. Extensive experiments demonstrate that NavCrafter achieves state-of-the-art novel-view synthesis under large viewpoint shifts and substantially improves 3D reconstruction fidelity.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!