3D 모델링과 공간 추론의 분리: 새로운 접근 방식
Disentangling 3D Modeling from Spatial Reasoning
본 연구에서는 대규모 학습을 통해 암묵적인 3차원 인식 및 추론을 동시에 습득하는 기존 방식 대신, 3차원 인식을 명시적으로 분리하여 공간 추론에 대한 대체 패러다임을 탐구합니다. 핵심적인 관찰은 현대의 인식 모델이 연속적인 3차원 기하학 정보를 정확하게 추정하는 데 뛰어난 반면, 대규모 언어 모델(LLM)은 특히 합성 및 상징적 추론에 효과적이라는 점입니다. 이러한 상호 보완적인 강점을 바탕으로, 우리는 기존의 전문적인 인식 모델을 사용하여 물리 세계를 구조화된 3차원 정보로 재구성하고, LoRA를 활용하여 LLM을 미세 조정하여 명시적인 기하학 정보만을 기반으로 추론하도록 하는 간단하면서도 효과적인 프레임워크인 Disentangled Spatial Reasoner (DiSR)를 제안합니다. DiSR은 대규모 3차원 질의응답 학습 데이터나 복잡한 도구 사용 정책 없이도, 기존의 공간 추론 벤치마크에서 경쟁력 있는 성능을 달성합니다. 뛰어난 성능 외에도, DiSR은 향상된 해석 가능성, 모듈성 및 계산 효율성을 제공하며, 이는 명시적인 인식과 추론의 분리가 공간 지능을 위한 엔드-투-엔드 모델링 방식에 대한 확장 가능하고 효과적인 대안임을 보여줍니다.
In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particularly effective at compositional and symbolic reasoning. Motivated by these complementary strengths, we propose the Disentangled Spatial Reasoner (DiSR), a simple yet effective framework that reconstructs the physical world into structured 3D evidence using off-the-shelf expert perception models and fine-tunes an LLM with LoRA to perform reasoning solely over this explicit geometric evidence. Without large-scale 3D VQA training or complex tool-use policies, DiSR achieves competitive performance on popular spatial reasoning benchmarks. Beyond its strong performance, DiSR offers improved interpretability, modularity, and computational efficiency, demonstrating that explicit separation of perception and reasoning is a scalable and effective alternative paradigm to end-to-end modeling for spatial intelligence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.