2606.11918v1 Jun 10, 2026 cs.AI

심문술의 기술: 일관성이 공간 추론에서의 사실성을 증폭시킨다

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

M. Ovsjanikov
M. Ovsjanikov
Citations: 11,148
h-index: 52
Théo Uscidda
Théo Uscidda
Citations: 192
h-index: 7
Leonidas J. Guibas
Leonidas J. Guibas
Citations: 7
h-index: 2
Marta Tintoré Gazulla
Marta Tintoré Gazulla
Citations: 8
h-index: 1
Federico Tombari
Federico Tombari
Citations: 18
h-index: 2

현재의 대규모 추론 모델(LRM)은 놀라운 일반적인 능력을 보이지만, 공간 추론 작업에서는 현저히 낮은 성능을 나타냅니다. 기존 접근 방식은 이러한 격차를 지식 부족으로 간주하며, 외부 시각 소스 또는 합성 엔진에서 가져온 레이블이 지정된 공간 데이터를 활용하여 지도 학습(SFT)을 통해 모델을 개선합니다. 이에 반해, 우리는 많은 작업에서 공간 추론 능력이 이미 사전 훈련된 LRM에 내재되어 있지만, 기하학적 2D 및 3D 제약 조건 하에서의 논리적 일관성을 통해 정렬될 필요가 있다고 주장합니다. 본 연구에서는 ground-truth 레이블 없이 내부 추론 과정을 목표로 하는 자기 지도 강화 학습(RL) 프레임워크를 제안합니다. 변환 하에서 기하학적 및 의미적 일관성을 검증하는 보상 함수인 '일관성 검증기'라는 개념을 공식화함으로써, 모델이 공간 추론 능력을 향상시킬 수 있음을 보여줍니다. 이미지 변환(예: 반전)과 텍스트 변환(예: 질문에서 객체의 순서 변경)을 모두 사용하며, 쌍별 검증기에 최적화된 새로운 그룹 상대 정책 최적화(Group Relative Policy Optimization)의 최소 매칭 변형인 OT-GRPO라는 새로운 최적 수송 기반 RL 전략을 제안합니다. 우리는 이러한 레이블이 없는 일관성 학습이 ground-truth 감독 학습으로 훈련된 모델의 정확도에 근접하며, 다양한 작업 및 데이터 도메인에서 유사한 일반화 성능을 달성한다는 것을 보여줍니다.

Original Abstract

Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled spatial data from external vision sources or synthetic engines. In contrast, we argue that for many tasks, spatial reasoning capabilities are already present in pre-trained LRMs but require alignment through logical coherence under geometric 2D and 3D constraints. In this work, we propose a self-supervised reinforcement learning (RL) framework that targets the internal reasoning process without requiring ground-truth annotations. By formalizing the notion of consistency verifiers -- reward functions that check for geometric and semantic consistency under transformations -- we demonstrate that models can improve their spatial reasoning abilities. We use both image transformations, like flipping, and textual transformations, like swapping the order of objects in the question, and propose a new optimal transport-based RL strategy, OT-GRPO, which is a minimal-matching variant of group relative policy optimization tailored to pairwise verifiers. We show that this label-free consistency training approaches the accuracy of models trained with ground-truth supervision and achieves similar generalization across diverse tasks and data domains.

0 Citations
0 Influential
26 Altmetric
130.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!