2601.13707v1 Jan 20, 2026 cs.CV

어텐션 공간 대비 학습 기반 가이드 방식으로 LVLM의 효율적인 환각 현상 완화

Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs

Yujin Jo
Yujin Jo
Citations: 4
h-index: 1
Sang-Peel Bae
Sang-Peel Bae
Citations: 67
h-index: 2
Taesup Kim
Taesup Kim
Citations: 24
h-index: 2

대규모 시각-언어 모델(LVLM)에서 발생하는 환각 현상은 종종 언어적 선입견이 시각적 증거보다 우세해져 객체 오인식 및 시각적으로 일관되지 않은 설명을 초래합니다. 본 연구에서는 환각 현상 완화를 대비 학습 기반 가이드 방식으로 접근하여, 시각적으로 근거가 있고 의미적으로 정확한 텍스트 생성을 유도합니다. 이러한 접근 방식은 모델의 내부 동작을 조절하여 언어적 선입견에 대한 과도한 의존성을 줄이고, 시각적으로 근거가 있는 표현과 언어만 사용한 표현을 대비시킵니다. 우리는 어텐션 공간 대비 학습 가이드(ACG)라는 단일 패스(single-pass) 메커니즘을 제안합니다. ACG는 자기 어텐션 레이어 내에서 작동하여, 단일 순방향 연산으로 시각-언어 및 언어만 사용한 어텐션 경로를 모두 생성합니다. 이러한 통합은 모델의 표현 맥락화에 직접적으로 내장되어 계산 효율적인 가이드를 가능하게 합니다. 단일 패스 방식에서 발생하는 근사 오차를 보정하기 위해, 언어만 사용한 경로와 정렬된 구성 요소를 제거하는 직교 보정을 추가적으로 적용하여 시각적 기여도를 선택적으로 증폭시킵니다. CHAIR 및 POPE 벤치마크에 대한 실험 결과, ACG는 최고 수준의 정확성과 캡션 품질을 달성하는 동시에 계산 비용을 크게 줄입니다. 본 연구는 기존의 다중 순방향 패스를 요구하는 대비 디코딩 방식에 비해 최대 2배까지 지연 시간을 줄이는, 원칙적이고 효율적인 대안을 제시합니다.

Original Abstract

Hallucinations in large vision-language models (LVLMs) often arise when language priors dominate over visual evidence, causing object misidentification and visually inconsistent descriptions. We address this issue by framing hallucination mitigation as contrastive guidance, steering generation toward visually grounded and semantically faithful text. This approach regulates the model's internal behavior by reducing over-dependence on language priors and contrasting visually grounded with language-only representations. We propose Attention-space Contrastive Guidance (ACG), a single-pass mechanism that operates within self-attention layers to construct both vision-language and language-only attention paths in a single forward computation. This integration enables computationally efficient guidance directly embedded in the model's representation contextualization. To correct approximation bias introduced by the single-pass formulation, we further apply an orthogonalized correction that removes components aligned with the language-only path, selectively amplifying visual contributions. Experiments on the CHAIR and POPE benchmarks show that ACG achieves state-of-the-art faithfulness and caption quality while significantly reducing computational cost. Our method establishes a principled and efficient alternative, reducing latency by up to 2x compared to prior contrastive decoding methods that require multiple forward passes.

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!