2605.27959v1 May 27, 2026 cs.CV

ROVER: 객체 중심 시각적 증거를 활용한 상황 인식 기반 멀티 이미지 추론

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning

Tingting Gao
Tingting Gao
Citations: 503
h-index: 12
Hongjian Dou
Hongjian Dou
Citations: 3
h-index: 1
Guannan Lv
Guannan Lv
Citations: 4
h-index: 1
R. Nie
R. Nie
Citations: 26
h-index: 1

최근 멀티모달 대규모 언어 모델(MLLM)은 심층적인 추론을 위해 시각적 정보를 더욱 세밀하게 위치시키고 통합하는 경향이 있습니다. 기존의 상황 인식 기반 접근 방식은 주로 관심 영역(RoI)에 초점을 맞추며, 잘라낸 이미지 패치나 RoI별 특징을 추론 과정에 주입합니다. 그러나 이러한 설계는 전체적인 장면 이해와 객체 간 관계를 약화시킬 수 있으며, RoI의 개수 및 크기에 따라 디코딩 비용이 증가하는 문제가 있습니다. 반면, 적응적 시각적 특징 선택 방법은 종종 세밀한 지도 학습이나 복잡한 휴리스틱을 필요로 합니다. 이러한 한계점을 해결하기 위해, 우리는 ROVER(Routing Object-centric Visual Evidence for grounded multi-image Reasoning)라는 가볍고 학습 가능한 플러그인을 제안합니다. ROVER는 각 객체 상황 인식 예측 시, 단계별 토큰 트리플렛을 주입하여 다음과 같은 효과를 얻습니다: (i) 현재 추론 컨텍스트를 통합하고, (ii) 객체 중심의 차등 주의(differential attention)를 통해 이미지 내 단서를 시각적 작업 공간으로 추출하며, (iii) 이 공간 내에서 객체 및 이미지 간의 과거 정보를 고려하여 증거를 전달하고 통합하여 후속 추론을 수행합니다. 우리는 ROVER를 Qwen2.5-VL-7B 모델에 통합하고, interleaved SFT-to-GRPO 학습 파이프라인을 개발했습니다. 원래 데이터셋과 평가 프로토콜을 엄격하게 준수하면서, 우리의 방법은 MM-GCoT에서 (+4.8% 답변 정확도, +14.6% 상황 인식 정확도) 및 VideoEspresso에서 (+8.6% 답변 정확도) 최고 성능을 달성했습니다. VideoEspresso로 학습된 모델은 다양한 벤치마크에서 평균적으로 기본 모델보다 +4.7% 더 높은 성능을 보여주며, 뛰어난 일반화 능력을 입증합니다.

Original Abstract

Multimodal Large Language Models (MLLMs) have increasingly localized and interleaved visual evidence for deliberative reasoning. Grounding-based approaches typically focus on regions of interest (RoIs) by injecting cropped image patches or RoI-specific features into the reasoning context. However, such designs can weaken holistic scene understanding and inter-object relations, while incurring decoding costs that scale with the number and size of RoIs. Alternatively, adaptive visual feature selection often requires fine-grained supervision or complex heuristics. To address these limitations, we propose ROVER (Routing Object-centric Visual Evidence for grounded multi-image Reasoning), a lightweight, learnable plugin for efficient global visual evidence routing. Upon each object grounding prediction, ROVER injects a step-specific token triplet to synergistically: (i) aggregate the ongoing reasoning context, (ii) distill intra-image cues into a visual working space via object-centric differential attention, and (iii) route and integrate history-aware evidence across objects and images within this space for subsequent reasoning. We integrate ROVER into Qwen2.5-VL-7B and develop an interleaved SFT-to-GRPO training pipeline. Strictly adhering to the original datasets and evaluation protocols, our method achieves the best performance on MM-GCoT (+4.8% answer accuracy, +14.6% grounding accuracy) and VideoEspresso (+8.6% answer accuracy). The VideoEspresso-trained model demonstrates strong transferability, outperforming the base model by +4.7% on average across diverse benchmarks.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!