2608.04759v1 Aug 05, 2026 cs.CV

추적, 검증 및 수정: 멀티모달 LLM의 공간 추론을 위한 학습이 필요 없는 프레임워크

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Zhaoxia Yin
Zhaoxia Yin
Citations: 100
h-index: 5

멀티모달 대규모 언어 모델(MLLM)은 상당한 발전을 이루었지만, 여전히 입력 이미지와 일치하지 않는 중간 판단을 내릴 수 있으며, 이는 추론 과정에서 오류를 발생시키고 최종 답변에 영향을 미칠 수 있습니다. 기존 방법들은 주로 학습 또는 추가적인 공간 정보를 통해 공간 추론 능력을 향상시키지만, 모델 입력에 대한 추론 과정 자체의 충실성을 고려하지 않습니다. 본 연구에서는 충실하지 않은 추론 체인이 최종 답변 정확도를 현저히 저하시킨다는 것을 보여줍니다. 이러한 문제를 해결하기 위해, 우리는 공간 추론 검증 및 수정에 사용될 수 있는 모듈화되고 학습이 필요 없는 프레임워크를 제안합니다. 이 프레임워크는 '공간 증거 그래프(Spatial Evidence Graph, SEG)'를 구축하여 체인-오브-생트(Chain-of-Thought) 추론에서 추출된 기본적인 공간 증거를 시각적 개체, 공간 관계, 출처 단계 및 시각적 증거와 연결합니다. '시각적 증거 신뢰도 평가(Spatial Evidence Reliability Assessment, SERA)'는 객체의 존재 여부, 위치 정보 및 기하학적 측정 값을 기반으로 시각적 증거의 신뢰도를 평가합니다. 그런 다음 프레임워크는 신뢰할 수 있는 시각적 증거에 의해 모순되는 가장 이른 시점의 공간 증거 단위를 식별하고, 원래 MLLM이 후속 추론과 최종 답변을 수정하도록 안내합니다. 15개의 모델-데이터셋 환경에서 본 방법은 평균 정확도 68.94%를 달성하여 비교 대상 기준 성능보다 평균 8.55%p 더 우수한 결과를 보였습니다. 본 연구의 코드는 공개될 예정입니다.

Original Abstract

Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image, allowing errors to propagate through the reasoning chain and affect the final answer. Existing methods mainly improve spatial reasoning through training or additional spatial information, without considering whether the reasoning process itself is faithful to the model input. Our study shows that unfaithful reasoning chains significantly reduce final-answer accuracy. To address this issue, we propose a modular and training-free framework for spatial reasoning verification and correction. The framework constructs a Spatial Evidence Graph (SEG), which associates atomic spatial evidence extracted from Chain-of-Thought reasoning with visual entities, spatial relations, source steps, and visual evidence. Spatial Evidence Reliability Assessment (SERA) evaluates the reliability of visual evidence based on object existence, localization, and geometric measurements. The framework then identifies the earliest spatial evidence unit contradicted by reliable visual evidence and guides the original MLLM to revise the subsequent reasoning and final answer. Across 15 model-dataset settings, our method achieves an average accuracy of 68.94%, outperforming the compared baselines by 8.55 percentage points on average. Our code will be open-sourced.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!