온 디맨드 시각 정보 활용: 다중 모드 추론을 위한 인지적 스케줄링 프레임워크
Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning
기존의 다중 모드 추론 방식은 주로 두 가지 패러다임을 따릅니다. 첫 번째는 시각 정보를 텍스트로 변환한 후 추론을 수행하는 것이고, 두 번째는 통일된 시각-언어 표현 공간 내에서 엔드 투 엔드 추론을 수행하는 것입니다. 이러한 방식들은 실증적인 성과를 거두었지만, 근본적인 구조적 한계를 가지고 있습니다. 첫 번째 방식은 정적인 시각-텍스트 변환에 의존하여 세부적인 시각 정보를 압축하고 손실시킬 수 있습니다. 두 번째 방식은 공동 최적화 및 어텐션 메커니즘으로 인해 언어 정보가 지배적인 역할을 하여, 추론 과정에서 시각 증거에 대한 충실도가 저하될 수 있습니다. 본 연구에서는 핵심적인 과제가 시각 정보를 추론 과정에 어떻게, 언제 도입하는가에 있다는 점을 강조합니다. 이러한 통찰력을 바탕으로, 우리는 CSMR이라는 다중 모드 추론 프레임워크를 제안합니다. CSMR은 언어 모델이 추론 과정을 제어하며, 작업과 관련된 중요한 시각 정보를 얻기 위해 독립적인 시각 인식 모듈을 언제 활성화할지 결정합니다. 여러 다중 모드 추론 벤치마크에서의 실험 결과는 CSMR이 대표적인 기본 방법보다 정확도 측면에서 일관되게 우수한 성능을 보임을 보여줍니다. 추가적인 실험 분석은 이러한 장점이 제안된 인지적 스케줄링 메커니즘으로부터 주로 비롯된다는 것을 확인합니다.
Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision-language representation space. Despite their empirical progress, both paradigms suffer from fundamental structural limitations. The former relies on static visual-to-text conversion, which tends to compress and lose fine-grained visual details. The latter is prone to linguistic dominance induced by joint optimization and attention mechanisms, leading to systematically weakened faithfulness to visual evidence during reasoning. In this work, we argue that a central challenge is how and when visual evidence is introduced into the reasoning process. Motivated by this insight, we propose CSMR, a multimodal reasoning framework in which a language model controls the reasoning process by deciding when to invoke an independent visual perception module to acquire task-relevant visual evidence. Experiments across multiple multimodal reasoning benchmarks show that CSMR consistently outperforms representative baseline methods in accuracy under a zero-shot setting. Further experimental analysis confirms that these advantages primarily arise from the proposed cognitive scheduling mechanism.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.