2607.24017v1 Jul 27, 2026 cs.CV

어텐션 공간에서 의미적 어텐션과 구조적 편향 분리 연구

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

Pengkun Jiao
Pengkun Jiao
Citations: 48
h-index: 4
Bin Zhu
Bin Zhu
Citations: 105
h-index: 6
Jingjing Chen
Jingjing Chen
Citations: 6,047
h-index: 37
Yu-Gang Jiang
Yu-Gang Jiang
Citations: 3,181
h-index: 27

멀티모달 대규모 언어 모델(MLLM)에서 어텐션 메커니즘의 성공적인 활용은 종종 그 내재된 미묘한 결함을 가립니다. 특히, MLLM은 일관되게 특정 의미적으로 유용한 시각 정보가 담긴 토큰에 비해 다른 시각 토큰에 과도하게 집중하는 경향을 보입니다. 이를 '레지스터(register)' 또는 '시각 어텐션 싱크(Visual Attention Sinks)'라고 부릅니다. 기존의 추론 개입 방법은 이러한 싱크 토큰을 식별하고 해당 어텐션 가중치를 재분배하려고 시도하지만, 이러한 접근 방식은 일반적으로 이러한 토큰을 개별적으로 처리하며 계산 효율성이 떨어지는 단점이 있습니다. 우리는 이 현상을 고립된 싱크 토큰을 넘어 광범위하게 작용하는 시각 특징에 대한 일반적인 텍스트 편향으로 재해석합니다. 이러한 관점에서, 만연한 구조적 편향은 의미 있는 시각 신호를 희석시키고, 모델이 유효한 시각 증거보다 언어적 선입견을 우선시하여 멀티모달 환각 현상을 초래하게 됩니다. 이러한 제한점을 해결하기 위해, 우리는 학습 과정 없이 적용 가능한 '살리언시 기반 정제 및 적응적 재분배(Saliency-guided Purification and Adaptive Redistribution, SPAR)'라는 개입 방법을 제안합니다. SPAR은 구조적 노이즈를 제거하고 확보된 어텐션 자원을 가장 유용한 시각 영역에 재분배하여 이러한 일반적인 텍스트 편향을 완화합니다. 다양한 환각 평가 지표에 대한 종합적인 실험 결과는 SPAR이 상당한 계산 오버헤드 없이 실제 시각 정보를 효과적으로 복원한다는 것을 보여줍니다.

Original Abstract

The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence. To address this limitation, we introduce Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention. SPAR mitigates this generalized textual bias by purifying structural noise and subsequently redistributing the reclaimed attention budget to the most informative visual regions. Comprehensive evaluations across a diverse spectrum of hallucination benchmarks demonstrate that SPAR effectively restores authentic visual grounding with negligible computational overhead.

0 Citations
0 Influential
18.5 Altmetric
92.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!