2605.28160v1 May 27, 2026 cs.AI

온 디맨드 시각 정보 활용: 다중 모드 추론을 위한 인지적 스케줄링 프레임워크

Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

Jiayi Ji
Jiayi Ji
Citations: 983
h-index: 16
Xiaoshuai Sun
Xiaoshuai Sun
Citations: 1,235
h-index: 20
Rongrong Ji
Rongrong Ji
Citations: 1,027
h-index: 17
Rui Zhao
Rui Zhao
Citations: 7
h-index: 1
Yidong Chen
Yidong Chen
Citations: 116
h-index: 5
Wujin Sun
Wujin Sun
Citations: 0
h-index: 0
Qianzhi Chen
Qianzhi Chen
Citations: 4
h-index: 1
Yang Zhang
Yang Zhang
Citations: 5
h-index: 2

기존의 다중 모드 추론 방식은 주로 두 가지 패러다임을 따릅니다. 첫 번째는 시각 정보를 텍스트로 변환한 후 추론을 수행하는 것이고, 두 번째는 통일된 시각-언어 표현 공간 내에서 엔드 투 엔드 추론을 수행하는 것입니다. 이러한 방식들은 실증적인 성과를 거두었지만, 근본적인 구조적 한계를 가지고 있습니다. 첫 번째 방식은 정적인 시각-텍스트 변환에 의존하여 세부적인 시각 정보를 압축하고 손실시킬 수 있습니다. 두 번째 방식은 공동 최적화 및 어텐션 메커니즘으로 인해 언어 정보가 지배적인 역할을 하여, 추론 과정에서 시각 증거에 대한 충실도가 저하될 수 있습니다. 본 연구에서는 핵심적인 과제가 시각 정보를 추론 과정에 어떻게, 언제 도입하는가에 있다는 점을 강조합니다. 이러한 통찰력을 바탕으로, 우리는 CSMR이라는 다중 모드 추론 프레임워크를 제안합니다. CSMR은 언어 모델이 추론 과정을 제어하며, 작업과 관련된 중요한 시각 정보를 얻기 위해 독립적인 시각 인식 모듈을 언제 활성화할지 결정합니다. 여러 다중 모드 추론 벤치마크에서의 실험 결과는 CSMR이 대표적인 기본 방법보다 정확도 측면에서 일관되게 우수한 성능을 보임을 보여줍니다. 추가적인 실험 분석은 이러한 장점이 제안된 인지적 스케줄링 메커니즘으로부터 주로 비롯된다는 것을 확인합니다.

Original Abstract

Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision-language representation space. Despite their empirical progress, both paradigms suffer from fundamental structural limitations. The former relies on static visual-to-text conversion, which tends to compress and lose fine-grained visual details. The latter is prone to linguistic dominance induced by joint optimization and attention mechanisms, leading to systematically weakened faithfulness to visual evidence during reasoning. In this work, we argue that a central challenge is how and when visual evidence is introduced into the reasoning process. Motivated by this insight, we propose CSMR, a multimodal reasoning framework in which a language model controls the reasoning process by deciding when to invoke an independent visual perception module to acquire task-relevant visual evidence. Experiments across multiple multimodal reasoning benchmarks show that CSMR consistently outperforms representative baseline methods in accuracy under a zero-shot setting. Further experimental analysis confirms that these advantages primarily arise from the proposed cognitive scheduling mechanism.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!