2607.27703v1 Jul 30, 2026 cs.AI

SpatialCLI: 공간 도구를 활용하여 추론하는 방법 학습, 그리고 그 없이도 추론하는 방법

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Sunzhu Li
Sunzhu Li
Citations: 54
h-index: 3
Shunyu Liu
Shunyu Liu
Citations: 646
h-index: 10
Shunian Chen
Shunian Chen
Citations: 1,534
h-index: 13
Yang Zhou
Yang Zhou
Citations: 0
h-index: 0
Zhuo Yang
Zhuo Yang
Citations: 0
h-index: 0
C.J. Yan
C.J. Yan
Citations: 0
h-index: 0
Jianyao Xu
Jianyao Xu
Citations: 17
h-index: 2
Weijie Fu
Weijie Fu
Citations: 0
h-index: 0
Peiliang Li
Peiliang Li
Citations: 5,510
h-index: 12
Yuxiang Cai
Yuxiang Cai
Citations: 0
h-index: 0
Zixuan Huang
Zixuan Huang
Citations: 0
h-index: 0
Chen Zhang
Chen Zhang
Citations: 0
h-index: 0
Xiaozhi Chen
Xiaozhi Chen
Citations: 7,423
h-index: 20

비전-언어 모델(VLMs)은 점점 더 많은 수의 에이전트에서 시각적 입력을 해석하고, 공간 관계에 대해 추론하며, 그러한 추론을 기반으로 작업 수준의 의사 결정을 내리는 데 사용됩니다. 그러나 근본적인 역량 불일치가 여전히 존재합니다. 일반적인 VLM은 전체 작업을 이해할 수 있지만 성공을 결정하는 시각적 세부 사항을 자주 놓치고, 반면 특수 비전 모델은 이러한 세부 사항을 포착할 수 있지만 이를 작업 수준의 의사 결정으로 변환할 수 없습니다. 본 연구에서는 SpatialCLI라는 프레임워크를 제안합니다. 이 프레임워크는 VLM이 공간 도구를 사용하여 추론하도록 훈련하고, 점진적으로 해당 도구가 제공하는 전문적인 인지 능력을 내재화하도록 합니다. SpatialCLI는 세 단계로 진행됩니다: (1) Call은 전문 비전 모델을 공간 도구로 활용하여 VLM의 인식을 향상시킵니다; (2) Learn은 콜드 스타트 SFT(Supervised Fine-Tuning)와 에이전트 기반 강화 학습을 사용하여 도구 사용 능력을 향상시킵니다; (3) Internalize는 성공적인 도구 사용 경로를 언어화하여 전문적인 인지 능력을 내재화합니다. 또한, SpatialCLI-Bench라는 516개의 예제로 구성된 벤치마크를 소개하며, 이는 위치 파악, 분할, 깊이 추정 및 자세 추정을 포함한 복합적인 인지에 대한 평가를 수행합니다. MindCube 데이터셋에서 SpatialCLI는 Qwen3-VL-8B-Instruct의 성능을 도구를 사용할 때 29.3%에서 84.6%로 향상시키며, 이는 도구를 사용하는 GPT-5.6 Sol (72.1%)보다 우수한 성능입니다. 또한, 내재화 과정을 거치면 도구가 없는 상태에서도 원래 성능인 73.8%를 유지합니다.

Original Abstract

Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!