2603.26348v1 Mar 27, 2026 cs.CV

정보 기반 검증을 통한 다중 모드 추론 능력 향상: 정보 획득 기반 검증

Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification

Chang Liu
Chang Liu
Citations: 37
h-index: 2
Feng Tang
Feng Tang
Citations: 3
h-index: 1
Yu-Jie Yuan
Yu-Jie Yuan
Citations: 149
h-index: 5
Aojun Zhou
Aojun Zhou
Citations: 1,738
h-index: 18
Xi Yang
Xi Yang
Citations: 91
h-index: 3
Yang Song
Yang Song
Citations: 27
h-index: 2
Shuai Lv
Shuai Lv
Citations: 13
h-index: 2
Kui Zhang
Kui Zhang
Citations: 19
h-index: 2

다중 모드 대규모 언어 모델(MLLM)은 뛰어난 다중 모드 추론 성능을 보이지만, 긴 형식의 생성 과정에서 반복적으로 나타나는 문제점이 있습니다. 모델이 생성하는 결과물이 길어질수록, 점차 이미지 증거에서 벗어나 텍스트 기반의 선입견에 의존하게 되어, 근거 없는 추론과 환각 현상이 발생합니다. 어텐션 분석 결과, MLLM은 시각 정보 검증 능력을 잠재적으로 가지고 있지만, 항상 활성화되지는 않는다는 것을 확인했습니다. 이러한 점에 착안하여, 우리는 추가적인 시각 입력 없이 MLLM이 스스로 시각적 검토를 수행할 수 있도록 하는 자기 진화형 훈련 프레임워크인 Visual Re-Examination (VRE)를 제안합니다. VRE는 더 강력한 모델로부터 시각적 능력을 전이하는 대신, 모델 자체를 활용하여 검토 기록을 생성함으로써, 모델 스스로 시각 정보를 활용하여 성능을 반복적으로 향상시킵니다. 다양한 다중 모드 벤치마크에서의 광범위한 실험 결과, VRE는 추론 정확도와 인지적 신뢰성을 지속적으로 향상시키고, 특히 긴 추론 과정에서 환각 현상을 크게 줄이는 것을 확인했습니다. 관련 코드는 https://github.com/Xiaobu-USTC/VRE 에서 확인할 수 있습니다.

Original Abstract

Multimodal Large Language Models (MLLMs) achieve strong multimodal reasoning performance, yet we identify a recurring failure mode in long-form generation: as outputs grow longer, models progressively drift away from image evidence and fall back on textual priors, resulting in ungrounded reasoning and hallucinations. Interestingly, Based on attention analysis, we find that MLLMs have a latent capability for late-stage visual verification that is present but not consistently activated. Motivated by this observation, we propose Visual Re-Examination (VRE), a self-evolving training framework that enables MLLMs to autonomously perform visual introspection during reasoning without additional visual inputs. Rather than distilling visual capabilities from a stronger teacher, VRE promotes iterative self-improvement by leveraging the model itself to generate reflection traces, making visual information actionable through information gain. Extensive experiments across diverse multimodal benchmarks demonstrate that VRE consistently improves reasoning accuracy and perceptual reliability, while substantially reducing hallucinations, especially in long-chain settings. Code is available at https://github.com/Xiaobu-USTC/VRE.

0 Citations
0 Influential
39.397207708399 Altmetric
0.0 Score
Original PDF
7

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!