2602.08241v1 Feb 09, 2026 cs.AI

MLLM은 실제로 보고 있는가: 멀티모달 LLM의 시각적 주의 강화

Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs

Siqu Ou
Siqu Ou
Citations: 2
h-index: 1
Tianrui Wan
Tianrui Wan
Citations: 13
h-index: 1
Xuelong Li
Xuelong Li
Citations: 220
h-index: 9
Zhiyuan Zhao
Zhiyuan Zhao
Citations: 350
h-index: 10
Junyu Gao
Junyu Gao
Citations: 728
h-index: 14

생각의 사슬(CoT) 추론은 복잡한 추론 과제에서 멀티모달 대형 언어 모델(MLLM)의 성능을 크게 향상시켰지만, 기존 접근 방식들은 주로 긴 텍스트 추론 경로에 의존하며 안정적인 시각적 주의(visual attention) 정책을 학습하는 기제는 제한적입니다. 본 연구의 분석에 따르면 현재의 MLLM은 시각적 초점이 약한 것으로 나타났습니다. 즉, 초기 단계의 시각적 불일치(misalignment)가 후속 추론 과정에서 거의 수정되지 않아 오류 전파 및 추론 실패로 이어집니다. 우리는 이러한 한계가 훈련 중 시각적 주의에 대한 불충분한 기여도 할당(credit assignment)에서 비롯된다고 주장합니다. 이 문제를 해결하기 위해, 우리는 영역(region) 수준의 시각적 주의 기반 보상을 도입한 강화 학습(RL) 프레임워크로 훈련된 시각적 추론 모델인 SAYO를 제안합니다. 이 보상은 최적화 신호를 시각적 근거가 있는 추론 단계와 명시적으로 정렬하여 모델이 더 신뢰할 수 있는 주의 행동을 학습하도록 합니다. 다수의 멀티모달 벤치마크에 대한 광범위한 실험 결과, SAYO는 다양한 추론 및 인식 과제에서 일관되게 성능을 향상시키는 것으로 나타났습니다.

Original Abstract

While chain-of-thought (CoT) reasoning has substantially improved multimodal large language models (MLLMs) on complex reasoning tasks, existing approaches largely rely on long textual reasoning trajectories and provide limited mechanisms for learning stable visual attention policies. Our analysis shows that current MLLMs exhibit weak visual focus: early-stage visual misalignment is rarely corrected during subsequent reasoning, leading to error propagation and failed inferences. We argue that this limitation stems from inadequate credit assignment for visual attention during training. To address this issue, we propose SAYO, a visual reasoning model trained with a reinforcement learning (RL) framework that introduces a region-level visual attention-based reward. This reward explicitly aligns optimization signals with visually grounded reasoning steps, enabling the model to learn more reliable attention behaviors. Extensive experiments across multiple multimodal benchmarks demonstrate that SAYO consistently improves performance on diverse reasoning and perception tasks.

2 Citations
1 Influential
7 Altmetric
39.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!