2606.12830v1 Jun 11, 2026 cs.CV

인식, 상호작용, 추론: 공간 추론을 위한 도구 기반 시각 에이전트 개발

Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning

Meng Lu
Meng Lu
Citations: 123
h-index: 2
Changye Li
Changye Li
Citations: 97
h-index: 5
Ligeng Zhu
Ligeng Zhu
Citations: 8,492
h-index: 23
Yi Wu
Yi Wu
Citations: 304
h-index: 2

최근의 비전-언어 모델(VLMs)은 강력한 다중 모드 이해 능력을 보여주지만, 능동적인 증거 획득과 다단계 시각 상호작용을 요구하는 공간 추론 작업에서는 여전히 한계를 드러냅니다. 이러한 제한 사항은 비전 인코더로부터 얻는 암시적 시각 표현만으로는 세밀한 공간 정보를 복원하기에 충분하지 않다는 것을 시사합니다. 본 연구에서는 지도 기반 공간 추론 작업을 위한 도구 기반 시각 에이전트인 PERIA(PERception-Interaction-reason Agent)를 소개합니다. PERIA는 텍스트, 기호 및 공간 증거를 노출하는 비전 인식 도구와 시각적 컨텍스트 조작, 경로 추적 및 공간 관계 검증을 위한 비전 상호 작용 도구라는 두 가지 경량화된 도구 패밀리를 사용합니다. PERIA의 학습을 위해, 우리는 지도 기반 도구 사용 트래jectory 합성, 복합 보상 및 Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO)을 결합한 통합 학습 방법을 개발했습니다. 8개의 데이터셋에서 추출된 13개의 벤치마크에 대한 실험 결과, PERIA-8B는 동일한 크기의 기존 최고 성능 모델보다 in-distribution 벤치마크에서 10.0%, out-of-distribution 벤치마크에서 4.4% 향상되었으며, Qwen3-VL-235B-A22B-Thinking 및 GPT-5와 같은 훨씬 더 큰 모델과 유사한 성능을 달성했습니다. 이는 PERIA가 공간 추론 능력을 향상시키는 데 효과적임을 입증합니다.

Original Abstract

While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limitation suggests that relying solely on implicit visual representations from vision encoders is insufficient for recovering fine-grained spatial evidence. We introduce PERception-Interaction-reason Agent (PERIA), a tool-augmented visual agent for spatial reasoning tasks across map reasoning, visual probing, and vision reconstruction. PERIA uses two lightweight tool families: vision perception tools for exposing textual, symbolic, and spatial evidence, and vision interaction tools for manipulating visual context, tracing paths, and verifying spatial relations. To train PERIA, we develop a unified recipe that combines supervised tool-use trajectory synthesis, composite rewards, and Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) for effective multi-tool behavior. Experiments on 13 benchmarks from 8 datasets show that PERIA-8B improves over the Qwen3-8B backbone by 10.0% on in-distribution benchmarks and 4.4% on out-of-distribution benchmarks, while outperforming previous state-of-the-art baselines of similar size by 7.0%-14.8%. It also achieves performance comparable to much larger models such as Qwen3-VL-235B-A22B-Thinking and GPT-5, demonstrating the effectiveness of PERIA in enhancing spatial reasoning capabilities.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!