2606.17678v1 Jun 16, 2026 cs.CV

먼저 보고 나중에 답변하기: 충분성 기반 강화 학습을 통한 시각 증거 사전 정렬

See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL

Guoshun Nan
Guoshun Nan
Citations: 75
h-index: 3
Junyi Zhu
Junyi Zhu
Citations: 35
h-index: 3
Yilian Liu
Yilian Liu
Citations: 162
h-index: 8
Sicong Leng
Sicong Leng
Citations: 2,536
h-index: 14
Jiayu Huang
Jiayu Huang
Citations: 0
h-index: 0
Xuancheng Zhu
Xuancheng Zhu
Citations: 0
h-index: 0
Yisong Chen
Yisong Chen
Citations: 0
h-index: 0
Zexian Wei
Zexian Wei
Citations: 0
h-index: 0
Xiaofeng Tao
Xiaofeng Tao
Citations: 220
h-index: 4
Ming Sun
Ming Sun
Citations: 0
h-index: 0

멀티모달 대규모 언어 모델(MLLM)은 강력한 텍스트 추론 능력과 시각적 입력을 통합하지만, 응답이 때때로 해당 이미지와 일치하지 않아 추론 과정에서 시각적 증거가 효과적으로 활용되지 않고 있음을 나타냅니다. 기존의 학습 방식은 일반적인 정렬을 위해 대규모 캡션 기반 사전 학습을 사용한 후, 지시사항 준수 및 복잡한 추론 능력을 향상시키기 위해 지도 학습 및 강화 학습을 수행합니다. 그러나 이러한 사전 학습은 상대적으로 약한 시각적 정보 연결을 제공하며, 짧고 일반적인 캡션은 모델이 중요한 객체에 편향되도록 유도하고 세부적인 시각적 증거는 무시하게 만듭니다. 본 논문에서는 사전 학습과 사후 학습 사이의 중간 단계인 Visual Evidence Pre-Alignment (VEPA)를 소개합니다. VEPA는 Group Relative Policy Optimization (GRPO)을 사용하여 질문에 조건화된 시각적 증거 설명을 최적화하는 새로운 충분성 기반 목표 함수를 탐색합니다. 다양한 벤치마크에서 수행한 광범위한 실험 결과, VEPA는 시각적으로 어려운 평가에서 일관되게 성능을 향상시키며 표준적인 지도 학습 사후 훈련과 상호 보완적인 효과를 나타냅니다. 추가 분석 결과, 이러한 성능 향상은 추가적인 작업별 학습이 아닌 강화된 전이 가능한 시각적 정보 연결 능력에서 비롯되는 것으로 확인되었습니다.

Original Abstract

Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffective utilization of visual evidence during inference. The prevailing training paradigm relies on large-scale caption-based pretraining for general alignment, followed by supervised fine-tuning and reinforcement learning to enable instruction following and complex reasoning. However, such pretraining provides only weak visual grounding: short, coarse captions bias models toward salient objects while neglecting fine-grained visual evidence. In this paper, we introduce Visual Evidence Pre-Alignment (VEPA), an intermediate stage between pretraining and post-training that explores a novel sufficiency-driven objective with Group Relative Policy Optimization (GRPO) to optimize question-conditioned visual evidence descriptions. Extensive experiments across diverse benchmarks show that our VEPA consistently enhances performance on visually demanding evaluations and complements standard supervised post-training. Further analyses show that the income stems from strengthened, transferable visual grounding, rather than from additional task-specific training.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!