2607.21595v1 Jul 23, 2026 cs.CV

3D-Aware VLMs with Implicit and Explicit Geometries

Quanhao Qian
Quanhao Qian
Citations: 19
h-index: 3
Gongjie Zhang
Gongjie Zhang
Citations: 19
h-index: 3
Ran Xu
Ran Xu
Citations: 165
h-index: 4
Wenhao Li
Wenhao Li
Citations: 13
h-index: 2
Xueying Jiang
Xueying Jiang
Citations: 155
h-index: 6
Deli Zhao
Deli Zhao
Citations: 19
h-index: 3
Shijian Lu
Shijian Lu
Citations: 65
h-index: 4

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RGB-only design injects strong 3D inductive biases for fine-grained spatial understanding and reasoning without requiring any additional 3D inputs. Extensive experiments show that VLM-IE3D achieves superior performance consistently across various 3D tasks including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning. Code and models are available at https://github.com/Vegetebird/VLM-IE3D.

0 Citations
0 Influential
36.540251005511 Altmetric
182.7 Score
Original PDF
14

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!