2607.29412v1 Jul 31, 2026 cs.CV

어텐션 헤드의 역할 파괴: VLM의 환각 현상 이해 및 탐지

Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs

Nan Duan
Nan Duan
Citations: 43
h-index: 3
Haoyang Huang
Haoyang Huang
Citations: 397
h-index: 5
Wenbo Li
Wenbo Li
Citations: 259
h-index: 2
Mingyu Wang
Mingyu Wang
Citations: 125
h-index: 5
Weilin Jin
Weilin Jin
Citations: 5
h-index: 2
Tong Jia
Tong Jia
Citations: 318
h-index: 9
Chaoran Luo
Chaoran Luo
Citations: 0
h-index: 0
Ying Li
Ying Li
Citations: 0
h-index: 0

시각-언어 생성 분야에서 상당한 발전이 있었음에도 불구하고, 비전-언어 모델(VLM)은 여전히 환각 현상을 일으키며 입력 이미지와 일치하지 않거나 지원되지 않는 내용을 생성하는 경향이 있습니다. 기존 연구는 주로 시각-텍스트 불균형과 같은 특정 유형의 환각 패턴을 중심으로 탐지 또는 완화 방법을 설계했지만, 실제 VLM 환각은 여러 패턴의 혼합에서 발생하며, 단일 패턴에 국한된 신호는 모델 및 작업 전반에 걸쳐 안정적으로 유지되기 어렵습니다. 우리는 통일된 헤드 레벨 관점에서 분석한 결과, 환각으로 인한 변화가 각 헤드의 충실한 문맥적 행동에서의 국소적인 편차로 나타나는 것을 발견했으며, 이를 '역할 파괴(Role-Break)'라고 명명했습니다. 자세한 분석에 따르면 이러한 편차는 어텐션 헤드, 문맥 소스 및 편차 방향에 따라 체계적으로 구성되어 있으며, 헤드 식별 정보가 유지되면 생성된 신호를 선형적으로 읽을 수 있습니다. 이러한 발견을 바탕으로, VLM의 미세 조정 없이 작동하는 가벼운 선형 탐지기를 개발했으며, 이 탐지기의 특징 차원은 5,000 이하이며, 6개의 VLM과 4개의 벤치마크에서 평균 AUROC가 93.23%를 달성했습니다. 소규모 개입 실험 결과, 탐지된 토큰은 판별적 설정에서 직접적으로 활용될 수 있는 것으로 나타났습니다.

Original Abstract

Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is inconsistent with or unsupported by the input image. Existing works largely design detection or mitigation methods around one specific hallucination pattern, such as visual-textual imbalance, but real VLM hallucinations arise from a mixture of multiple patterns, so signals bound to a single pattern struggle to remain stable across models and tasks. Under a unified head-level view, we find that hallucination-induced changes manifest as localized deviations from each head's faithful contextual behavior, a phenomenon we term Role-Break. Detailed analysis reveals that these deviations are systematically organized across attention heads, contextual sources, and deviation directions, and that the resulting signal is linearly readable once head identity is preserved. Based on these findings, we build a lightweight linear detector on top of Role-Break that requires no fine-tuning of the VLM, whose feature dimension stays below 5,000 and reaches an average AUROC of 93.23 across six VLMs and four benchmarks. A small-scale intervention experiment further shows that the detected tokens can be directly acted upon in the discriminative setting.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!