2608.03450v1 Aug 04, 2026 cs.MM

효율성과 효과의 균형: 명시적 및 잠재적 사고 간의 훈련 없이 어텐션 가이드 방식으로 전환하는 다중 모드 대규모 언어 모델

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs

Haoqiang Kang
Haoqiang Kang
Citations: 436
h-index: 10
Jinpeng Wang
Jinpeng Wang
Citations: 25
h-index: 2
Bin Chen
Bin Chen
Citations: 97
h-index: 6
Yaowei Wang
Yaowei Wang
Citations: 141
h-index: 6
Ke Chen
Ke Chen
Citations: 20
h-index: 3
Kuofeng Gao
Kuofeng Gao
Citations: 632
h-index: 14
Zhenyu Lu
Zhenyu Lu
Citations: 7
h-index: 1
Liupeng Li
Liupeng Li
Citations: 10
h-index: 2

다중 모드 대규모 언어 모델(MLLM)에서의 추론은 세밀한 시각 인식과 엄격한 논리적 추론을 모두 필요로 합니다. 명시적인 텍스트 기반의 체인 오브 소트(CoT)는 계산 비용이 많이 들고 시각적 환각에 취약하며, 기존의 잠재적 추론 방법은 일반적으로 상당한 학습 비용이 필요합니다. 또한, 훈련 없이 LLM 추론 메커니즘을 다중 모드 환경으로 직접 적용하면 성능이 불안정해지는 경향이 있습니다. 이러한 실패는 토큰 수준의 엔트로피에 대한 의존성에서 비롯되며, 이는 근본적으로 인식적 불확실성(예: 명확하지 않은 시각적 세부 사항)과 논리적 불확실성(예: 복잡한 추론 단계)을 혼동합니다. 이러한 병목 현상을 극복하기 위해, 우리는 다중 모드 MLLM에 대한 새로운 훈련이 필요 없는 추론 전략을 제시하며, 이를 통해 인식 및 추론을 명시적으로 분리합니다. 제안하는 프레임워크인 어텐션 가이드 스위칭(AGS)은 모델의 인지적 집중도를 동적으로 측정하기 위한 새로운 지표인 '비전-텍스트 어텐션 비율'을 사용합니다. 이 지표에 따라, AGS는 연속적인 공간에서 고품질 시각 정보를 유지하기 위해 인식 관련 토큰에 대해 잠재적 추론을 적응적으로 활성화하고, 동시에 구조적 안정성을 유지하기 위해 논리 관련 토큰에 대해서는 명시적인 텍스트 생성을 강제합니다. 광범위한 실험 결과, 제안하는 방법은 최첨단 성능을 달성하며, 자기 회귀 단계 및 지연 시간을 줄여 정확도와 추론 효율성을 크게 향상시키는 것을 보여줍니다. 코드 배포 위치: https://github.com/swordAndSnow/MM26-AGS.

Original Abstract

Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based Chain-of-Thought (CoT) is computationally expensive and prone to visual hallucinations, while existing latent reasoning methods typically require costly training. Furthermore, directly adapting training-free LLM reasoning mechanisms to the multimodal setting yields unstable performance. We identify that this failure stems from their reliance on token-level entropy, which fundamentally conflates perceptual ambiguity (e.g., unclear visual details) with logical uncertainty (e.g., complex reasoning steps). To overcome this bottleneck, we present a novel training-free inference strategy for MLLMs that explicitly decouples perception and reasoning. We propose a novel metric, the vision-to-text attention ratio, to dynamically gauge the model's cognitive focus. Guided by this metric, our proposed framework, Attention-Guided Switching (AGS), adaptively triggers latent reasoning for perceptual tokens to preserve high-fidelity visual information in the continuous space, while enforcing explicit text generation for logical tokens to maintain structural anchoring. Extensive experiments demonstrate that our method achieves state-of-the-art performance, significantly improving both accuracy and inference efficiency by reducing autoregressive steps and latency. Code is released at https://github.com/swordAndSnow/MM26-AGS.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!