이유를 따지고 다시 생각해보기: 다중 시점 재검토가 공간 추론 성능을 향상시킨다
Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning
일반적인 비디오에서 공간 추론은 카메라의 움직임에 의해 관찰 가능한 정보가 제한되기 때문에 본질적으로 어려운 문제입니다. 기존 방법들은 단일 단계 추론에 의존하며, 모델이 검증 가능한 증거 대신 의미론적 선행 지식을 통해 기하학적 모호성을 해결하도록 강제합니다. 우리는 공간 추론은 재검토 가능해야 한다고 주장합니다: 제한된 정보 하에서 형성된 결론은 상호 보완적인 시점이 제공될 때 수정될 수 있어야 합니다. 이러한 통찰력을 바탕으로, 저희는 학습이 필요 없는 추론 시간 프레임워크인 Reason, then Re-reason (ReRe)를 제안합니다. 이 프레임워크는 두 단계로 구성됩니다: 첫 번째 단계(Reason Phase)에서 MLLM은 원래 비디오로부터 공간 가설을 형성하고, 두 번째 단계(Re-reason Phase)에서는 합성된 새로운 시점의 비디오를 관찰하여 가설을 검증하거나 수정합니다. 효과적인 다중 시점 재검토를 가능하게 하기 위해, 저희는 예측된 3차원 기하 정보를 기반으로 전략적으로 상호 보완적인 새로운 시점을 생성하는 Geometry-to-Video 파이프라인을 설계했습니다. 이러한 시점은 장면 전체를 포괄하는 고각 및 비스듬한 관점을 제공하며, MLLM의 기존 비디오 인터페이스를 변경하지 않고 유지합니다. VSI-Bench와 STI-Bench에 대한 광범위한 평가 결과, ReRe는 오픈 소스 MLLM의 성능을 크게 향상시켜 독점적인 최첨단 수준에 근접하는 것을 보여주었습니다. 프로젝트 페이지: https://zhenjiemao.github.io/ReRe/
Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single-turn inference, forcing models to resolve geometric ambiguity through semantic priors rather than verifiable evidence. We argue that spatial reasoning should be revisitable: conclusions formed under limited evidence should remain open to revision when complementary viewpoints become available. Building on this insight, we propose Reason, then Re-reason (ReRe), a training-free, inference-time framework with two phases: in the Reason Phase, an MLLM forms a spatial hypothesis from the original video; in the Re-reason Phase, it verifies or revises the hypothesis by observing a synthesized novel-view video. To enable effective cross-view revisiting, we design a Geometry-to-Video pipeline that renders strategically complementary novel views from predicted 3D geometry. These views feature an elevated, oblique perspective with scene-spanning coverage, while preserving the MLLM's native video interface without architectural modifications. Extensive evaluations on VSI-Bench and STI-Bench demonstrate that ReRe substantially boosts open-source MLLMs to rival proprietary state-of-the-art performance. Project page: https://zhenjiemao.github.io/ReRe/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.