2601.22228v1 Jan 29, 2026 cs.CV

공간 속에서 길을 잃다? 비전-언어 모델이 상대 카메라 자세 추정에서 겪는 어려움

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation

Yftah Ziser
Yftah Ziser
Citations: 760
h-index: 15
Yifu Qiu
Yifu Qiu
University of Edinburgh;Cambridge University
Citations: 242
h-index: 8
Shay B. Cohen
Shay B. Cohen
Citations: 236
h-index: 9
Ken Deng
Ken Deng
Citations: 20
h-index: 3
Yoni Kasten
Yoni Kasten
Citations: 3,104
h-index: 12

비전-언어 모델(VLM)은 2차원 인식 및 의미 추론에서 뛰어난 성능을 보이지만, 3차원 공간 구조에 대한 이해는 제한적입니다. 본 연구에서는 상대 카메라 자세 추정(RCPE)이라는 기본적인 시각적 과제를 통해 이러한 간극을 조사합니다. RCPE는 두 이미지 쌍으로부터 상대적인 카메라의 이동 및 회전을 추론하는 작업입니다. 우리는 비언어적 주석이 포함된 비표시된 개인 시점 동영상에서 파생된 벤치마크인 VRRPI-Bench를 소개합니다. VRRPI-Bench는 공통 객체 주변의 동시 이동 및 회전을 반영하는 현실적인 시나리오를 모델링합니다. 또한, 개별적인 운동 자유도를 분리하는 진단 벤치마크인 VRRPI-Diag를 제안합니다. RCPE는 비교적 간단한 작업이지만, 대부분의 VLM은 얕은 2차원 휴리스틱을 벗어나 일반화하지 못하며, 특히 깊이 변화 및 광학 축을 따른 회전 변환에서 어려움을 겪습니다. 최첨단 모델인 GPT-5조차도 (정확도 0.64) 기존의 기하학적 기준 모델 (정확도 0.97) 및 인간 수준의 성능 (정확도 0.92)에 미치지 못합니다. 또한, VLM은 프레임 간의 공간적 단서를 통합할 때 일관성 없는 성능(최고 59.7%)을 보이는 등, 다중 이미지 추론에서도 어려움을 나타냅니다. 이러한 결과는 VLM이 3차원 및 다중 시점 공간 추론에 대한 이해가 부족하다는 것을 시사합니다.

Original Abstract

Vision-Language Models (VLMs) perform well in 2D perception and semantic reasoning compared to their limited understanding of 3D spatial structure. We investigate this gap using relative camera pose estimation (RCPE), a fundamental vision task that requires inferring relative camera translation and rotation from a pair of images. We introduce VRRPI-Bench, a benchmark derived from unlabeled egocentric videos with verbalized annotations of relative camera motion, reflecting realistic scenarios with simultaneous translation and rotation around a shared object. We further propose VRRPI-Diag, a diagnostic benchmark that isolates individual motion degrees of freedom. Despite the simplicity of RCPE, most VLMs fail to generalize beyond shallow 2D heuristics, particularly for depth changes and roll transformations along the optical axis. Even state-of-the-art models such as GPT-5 ($0.64$) fall short of classic geometric baselines ($0.97$) and human performance ($0.92$). Moreover, VLMs exhibit difficulty in multi-image reasoning, with inconsistent performance (best $59.7\%$) when integrating spatial cues across frames. Our findings reveal limitations in grounding VLMs in 3D and multi-view spatial reasoning.

1 Citations
0 Influential
7.5 Altmetric
38.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!