2608.07012v1 Aug 07, 2026 cs.CV

Scenix: 실행 가능한 장면 프로그램을 이용한 희소 시점 기반 3D 장면 복원

Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

Keyang Luo
Keyang Luo
Citations: 6
h-index: 2
Xin Wang
Xin Wang
Citations: 160
h-index: 4
Weikai Chen
Weikai Chen
Citations: 657
h-index: 13
Kai Li
Kai Li
Citations: 81
h-index: 5
Jierui Zhang
Jierui Zhang
Citations: 0
h-index: 0
Yingda Yin
Yingda Yin
Citations: 21
h-index: 3
Runze Zhang
Runze Zhang
Citations: 136
h-index: 5
Zhenyang Li
Zhenyang Li
Citations: 7
h-index: 1
Lutao Jiang
Lutao Jiang
Citations: 33
h-index: 3
Jiayu Dong
Jiayu Dong
Citations: 210
h-index: 7
Kai Yan
Kai Yan
Citations: 95
h-index: 4
Xiaoyang Huang
Xiaoyang Huang
Citations: 0
h-index: 0
Xiang Zhao
Xiang Zhao
Citations: 0
h-index: 0

적은 수의 보정되지 않은 RGB 이미지로부터 구조화되고 편집 가능한 3D 실내 장면을 생성하는 것은 고품질 개별 자산 생성 이상을 필요로 합니다. 시스템은 방의 구조를 추론하고, 불완전한 관찰에서 객체를 연결하며, 전체적으로 일관된 공간 구성을 복원해야 합니다. 기존 방법들은 주로 텍스트 입력 기반의 3D 장면 생성에 초점을 맞추거나, 연속적인 시각 정보와 추가적인 사전 지식(예: 인간이 주석을 단 마스크 또는 정확한 3D 레이아웃)을 필요로 하며, 이는 많은 노동력을 요구하고 일반적인 경우에 적용하기 어렵습니다. 본 논문에서는 실행 가능한 장면 프로그램을 이용한 희소 시점 기반 3D 장면 복원 프레임워크인 extsc{Scenix}를 제시합니다. extsc{Scenix}는 희소 시점을 입력으로 받아, 인지 기반 자산 생성 및 폐루프 공간 정제를 통해 실행 가능한 장면 프로그램을 예측합니다. 본 연구에서는 약 11만 개의 합성 및 실제 실내 장면 데이터셋( extit{dataset})을 구축했습니다. 이 데이터셋은 멀티뷰 이미지, 방 구조, 객체 중심 설명 및 메트릭 공간 정보를 포함하고 있습니다. 또한, 입력 시각 정보와 목표 장면을 일치시키는 관찰 기반 지도 학습 방법을 도입하여 성능 향상을 도모했습니다. 실험 결과는 extsc{XScene} 데이터셋, 실제 실내 이미지, 그리고 SpatialGen 데이터셋에 대한 평가를 통해 구조화된 장면 예측, 객체 위치 파악 및 공간 정제의 능력을 검증합니다.

Original Abstract

Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!