2606.11719v1 Jun 10, 2026 cs.CV

Ouroboros-Spatial: 공간 추론을 위한 데이터-모델 루프 완성

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning

Di He
Di He
Citations: 513
h-index: 8
Yuanrui Zhang
Yuanrui Zhang
Citations: 4
h-index: 1
E. Zhao
E. Zhao
Citations: 0
h-index: 0
Wei Wu
Wei Wu
Citations: 201
h-index: 6
Xueliang Zhao
Xueliang Zhao
Citations: 274
h-index: 5

공간 추론은 다중 모드 대규모 언어 모델(MLLM)에 대한 지속적인 과제입니다. 기존 접근 방식은 대부분 방대한 규모의 정적 큐레이션 데이터 세트에 의존하며, 여기서 모든 학습 샘플이 모델의 변화하는 능력과 관계없이 동일하게 취급됩니다. 이러한 정적인 패러다임은 본질적으로 데이터 효율성이 낮습니다. 학습 용량은 종종 현재 단계에서 모델에게 너무 쉽거나 너무 어려운 샘플에 사용됩니다. 이 한계를 해결하기 위해, 우리는 Ouroboros-Spatial을 제안합니다. 이는 모델이 제안자(proposer)와 솔버(solver)라는 이중 역할을 수행하는 자기 진화 학습 프레임워크입니다. 각 반복에서, 고정된 제안자는 3D 장면 메타데이터 및 원시 비디오 프레임을 기반으로 공간 질의-응답(QA) 쌍을 생성하며, 신뢰할 수 있는 정답을 얻기 위한 실행 가능한 코드를 함께 제공합니다. 학습 가능한 솔버는 선택된 샘플에 대해 미세 조정되며, 각 샘플에 대한 예측 신뢰도는 난이도 신호로 사용됩니다. 이 신호는 다음 반복에서 제안자로 피드백되어 솔버의 현재 능력에 더 적합한 질문을 생성하도록 유도합니다. 이러한 폐쇄 루프 설계 덕분에 학습 분포가 모델 능력과 함께 공동으로 진화하여 불필요하게 쉬운 예제를 줄이고, 제한된 학습 가치를 갖는 모호하거나 정보가 부족한 샘플을 필터링합니다. Ouroboros-Spatial은 6개의 공간 추론 벤치마크에서 Qwen3-VL-4B 및 Qwen3-VL-8B 모델의 성능을 크게 향상시키면서, 최근 대규모 큐레이션 데이터 세트보다 훨씬 적은 수의 학습 예제를 사용합니다. VSI-Bench에서는 각각 4B 및 8B 모델에 대해 9.9점과 6.8점이라는 상당한 성능 향상을 보여주며, 두 모델 모두 다양한 강력한 오픈 소스 및 독점 기반 모델을 능가할 수 있게 합니다.

Original Abstract

Spatial reasoning remains a persistent challenge for multimodal large language models (MLLMs). Existing approaches largely rely on large-scale, statically curated datasets, where all training samples are treated uniformly regardless of the model's evolving capabilities. This static paradigm is inherently data-inefficient: training capacity is often spent on samples that are either trivial or overly difficult for the model at its current stage. To address this limitation, we propose Ouroboros-Spatial, a self-evolving training framework in which the model plays dual roles as a proposer and a solver. In each iteration, a frozen proposer generates spatial question-answer (QA) pairs from 3D scene metadata and raw video frames, together with executable code for deriving reliable ground truth. A learnable solver is then fine-tuned on the accepted samples, and its per-sample prediction confidence is used as a difficulty signal. This signal is fed back to the proposer in the next iteration, guiding it to generate questions better matched to the solver's current capabilities. Through this closed-loop design, the training distribution co-evolves with model ability, reducing redundant trivial examples while filtering out ambiguous or uninformative samples with limited learning value. Across six spatial reasoning benchmarks, Ouroboros-Spatial substantially improves Qwen3-VL-4B and Qwen3-VL-8B while using an order of magnitude fewer training examples than recent large-scale curated datasets. On VSI-Bench, it yields absolute gains of 9.9 and 6.8 points for the 4B and 8B models, respectively, enabling both to outperform a wide range of strong open-source and proprietary baselines.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!