2606.11770v1 Jun 10, 2026 cs.AI

SVoT: 강화 학습을 통한 공간 추론을 위한 상태 인식 시각적 사고

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning

N. Lipovetzky
N. Lipovetzky
Citations: 1,644
h-index: 20
Yanbei Jiang
Yanbei Jiang
Citations: 44
h-index: 4
Chao Lei
Chao Lei
Citations: 39
h-index: 3
Markus Hiller
Markus Hiller
Citations: 1,107
h-index: 10
Zhijian Zhou
Zhijian Zhou
Citations: 12
h-index: 2
Xunye Tian
Xunye Tian
Citations: 13
h-index: 2
Krista A. Ehinger
Krista A. Ehinger
Citations: 244
h-index: 6

다중 모드 대규모 언어 모델(MLLM)에서 공간 추론은 여전히 어려운 과제이며, 이는 중간 상태와 상태 전환에 대한 신뢰할 수 있는 다단계 추론이 필요하기 때문입니다. 기존 연구에서는 종종 중간 상태를 검증하지 않고 상태 전환을 암시적인 과정으로 취급하여 다단계 공간 추론의 신뢰성을 제한합니다. 이러한 문제를 해결하기 위해 우리는 상태 인식을 갖춘 시각적 사고(SVoT)라는 강화 학습 프레임워크를 제안합니다. SVoT는 얽히고설킨, 검증 가능한 중간 상태와 시각화를 생성하며, 전환 추론 체인을 생성 과정에 통합하여 모델이 텍스트 및 시각적 추론을 통해 행동의 전제 조건과 효과를 확인하도록 합니다. 우리는 그룹 상대 정책 최적화(GRPO)를 사용하여 SVoT를 학습시키고, 보상 설계 방식을 통해 검증을 구현하고 다양한 미세 조정된 보상의 효능을 평가합니다. 기존 벤치마크에서 상태 전환을 단일 변수 업데이트로 단순화하여 문제를 크게 간소화하는 점을 고려하여, 우리는 고전 환경을 확장하고 다중 객체 상호 작용과 수리 추론이 필요한 새로운 환경인 Pacman 및 Gather를 도입하여 총 다섯 가지 도메인을 구축했습니다. 이러한 도메인은 생성된 중간 상태와 전환 추론에 대한 정량적 검증을 통해 다단계 공간 추론을 체계적으로 평가할 수 있도록 지원합니다. 전환 인식을 갖춘 감독 하에서 학습된 SVoT는 제안된 모든 도메인에서 최첨단 성능을 달성했으며, 분산 데이터 세트에서 최대 65%의 절대 정확도 향상을 보였습니다.

Original Abstract

Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate states and state transitions. Current studies often leave intermediate states unverified and treat state transitions as implicit processes, which limits reliability in multi-hop spatial reasoning. To address this, we propose State-aware Visualization-of-Thought (SVoT), a reinforcement learning framework that generates interleaved, verifiable intermediate states and visualizations. SVoT integrates transition reasoning chains into the generation processes, enabling the model to verify action preconditions and effects through interleaved textual and visual reasoning. We train SVoT via Group Relative Policy Optimization (GRPO), instantiating verification through reward design and evaluating the efficacy of different fine-grained rewards. As existing benchmarks reduce state transitions to single-variable updates, substantially simplifying the problems, we establish five domains by extending classical environments and introducing two novel domains, Pacman and Gather, that require multi-object interactions and numerical reasoning. These domains support systematic evaluation of multi-hop spatial reasoning with quantitative verification of generated intermediate states and transition reasoning. SVoT with transition-aware supervision achieves state-of-the-art performance across the introduced domains, yielding up to a 65% absolute accuracy gain on out-of-distribution test sets.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!