텍스트 기반의 비-맨해튼 환경에서의 3차원 실내 장면 생성
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
대규모 언어 모델(LLM)은 맨해튼 환경에서 3차원 실내 장면을 생성하는 데 놀라운 능력을 보여주었습니다. 그러나 기존 방법들은 종종 비-맨해튼 환경에서 현실적인 객체 배치 패턴을 포착하지 못합니다. 이는 주로 직각이 아닌 공간 관계를 모델링하는 데 어려움을 겪기 때문에 발생하며, 이로 인해 기하학적 오류가 많아지고 물리적 타당성이 낮아집니다. 이러한 문제를 해결하기 위해, 우리는 복잡한 비-맨해튼 환경에서 물리적으로 타당한 실내 장면을 생성하도록 설계된 새로운 텍스트 기반 프레임워크인 SPG-Layout을 제안합니다. 구체적으로, 우리는 객체 분포에 대한 통계적 사전 정보를 활용하여 학습 과정을 안내하고, 환경 이해와 타당성을 향상시킵니다. 또한, 인간의 디자인 워크플로우를 반영하여 큰 객체의 배치 우선순위를 정하는 계층적 배치 전략을 채택함으로써, 배치 오류를 크게 줄입니다. 이러한 구성 요소들을 결합하여 SPG-Layout은 의미론적 현실성과 물리적 타당성의 균형 잡힌 최적화를 달성합니다. 이러한 복잡한 환경에서의 성능을 평가하기 위해, 우리는 500개의 다양한 비-맨해튼 환경으로 구성된 새로운 벤치마크를 구축했습니다. 광범위한 실험 결과는 SPG-Layout이 맨해튼 및 비-맨해튼 환경 모두에서 기존 방법보다 일관되고 현저하게 우수한 성능을 보임을 보여줍니다. 코드 공개 예정입니다.
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge, we propose SPG-Layout, a novel text-driven framework designed to generate physically plausible indoor scenes within complex non-Manhattan environments. Specifically, we first utilize statistical priors of object distributions to guide the training process, enhancing environmental understanding and fidelity. Furthermore, mirroring human design workflows, we adopt a hierarchical layout strategy that prioritizes the placement of large objects, thereby substantially minimizing layout violations. By synergizing these components, SPG-Layout achieves a balanced optimization of semantic realism and physical plausibility. To evaluate performance in these complex settings, we constructed a new benchmark comprising 500 diverse non-Manhattan environments. Extensive experiments demonstrate that SPG-Layout consistently and significantly outperforms existing methods across both Manhattan and non-Manhattan environments. The code will be publicly released.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.