2607.13454v1 Jul 15, 2026 cs.CV

GeoAnchor: 잠재 공간 분해를 통한 협력적 추론을 활용한 3차원 공간 이해

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Xin Wei
Xin Wei
Citations: 23
h-index: 1
Hao Li
Hao Li
Citations: 313
h-index: 5
Hongbo Sun
Hongbo Sun
Citations: 31
h-index: 3
Han Fang
Han Fang
Citations: 95
h-index: 5
Zixin Pan
Zixin Pan
Citations: 0
h-index: 0
Yu Yu
Yu Yu
Citations: 0
h-index: 0
Jinglin Xu
Jinglin Xu
Citations: 41
h-index: 3
Zhiyu Lin
Zhiyu Lin
Citations: 0
h-index: 0
Ye Yuan
Ye Yuan
Citations: 26
h-index: 2
Zhongjiang He
Zhongjiang He
Citations: 177
h-index: 7
Hao Sun
Hao Sun
Citations: 24
h-index: 1

다중 모달 대규모 언어 모델(MLLM)이 놀라운 발전을 이루었지만, 2D 이미지로부터 3차원 공간 관계를 이해하는 것은 여전히 중요한 과제입니다. 기존 방법은 주로 기호 텍스트 토큰에 의존하는데, 이는 연속적인 기하학적 정보를 표현하는 데 있어 본질적인 한계를 가지고 있습니다. 최근 연구에서는 잠재 표현을 사용하여 추론 능력을 향상시키지만, 단일 유형의 잠재 변수만으로는 다양한 공간 태스크에 적응할 수 없으며, 복잡한 기하학적 시나리오에서 오차가 발생할 수 있습니다. 이러한 제한 사항을 해결하기 위해, 본 연구에서는 텍스트와 잠재 표현을 결합하여 추론하는 프레임워크인 GeoAnchor를 제안합니다. GeoAnchor는 3차원 공간 정보를 세 가지 상호 보완적인 구성 요소로 분해합니다: 객체 위치 정보에 대한 위치 잠재 변수, 관계 방향에 대한 방향 잠재 변수, 그리고 장면 구조에 대한 기하학 잠재 변수입니다. 이러한 구성 요소들은 구조화된 공간에서 재조합되어 지역적 증거를 생성하고 동시에 전체적인 맥락을 파악하여 동적이고 해석 가능한 추론을 가능하게 합니다. 또한, 모델이 지역적인 공간 인지로부터 포괄적인 3차원 이해에 이르도록 안내하는 협력적 학습 전략을 도입했습니다. 다양한 복잡한 3차원 추론 태스크에 대한 광범위한 실험 결과는 GeoAnchor가 최첨단 기술보다 우수한 성능을 보임을 입증하며, 그 효과성과 일반화 능력을 검증합니다.

Original Abstract

Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent continuous geometric information. While recent methods use latent representations to enhance reasoning, relying on a single latent type cannot adapt to the diversity of spatial tasks, leading to misalignment in complex geometric scenarios. To address these limitations, we propose GeoAnchor, an interleaved text-latent reasoning framework. GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure. These components are recombined in a structured space to construct local evidence while capturing global context, enabling dynamic and interpretable reasoning. Furthermore, we introduce a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding. Extensive experiments on diverse and complex 3D reasoning tasks demonstrate that GeoAnchor outperforms the state of the art, validating its effectiveness and generalization capabilities.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!