2605.25377v1 May 25, 2026 cs.CV

적대적 직교 분리 기법을 이용한 LVLM 환각 완화

Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xingjun Ma
Xingjun Ma
Citations: 704
h-index: 15
Haoxuan Ma
Haoxuan Ma
Citations: 35
h-index: 4
Ranjie Duan
Ranjie Duan
Citations: 4
h-index: 1
Ruoxi Cheng
Ruoxi Cheng
Citations: 37
h-index: 3
Ziyi Ye
Ziyi Ye
Citations: 765
h-index: 8
Zhengfei Hai
Zhengfei Hai
Citations: 1
h-index: 1
Tianle Zhang
Tianle Zhang
Citations: 1
h-index: 1
Xu Yang
Xu Yang
Citations: 8
h-index: 2

대규모 시각-언어 모델(LVLM)은 다중 모드 이해 능력이 향상되었지만, 생성된 내용이 실제 시각 정보와 충돌하는 환각 현상으로 인해 신뢰성이 제한됩니다. 기존의 환각 완화 방법들은 비용이 많이 드는 외부 개입 방식(예: 지시 조정 및 검색)에 의존하거나, 결함 있는 어텐션 가중치와 얽힌 잠재 표현이라는 한계점을 가진 내부 메커니즘을 사용합니다. 본 연구에서는 LVLM의 환각 현상을 완화하기 위한 잠재 기하학적 프레임워크인 적대적 직교 분리(AOD)를 제안합니다. AOD는 최소-최대 객관 함수를 통해 환각과 관련된 방향을 학습합니다. 분류기는 환각 신호를 투영된 구성 요소에 집중시키는 반면, 적대 네트워크는 그래디언트 역전파 레이어를 사용하여 직교 잔여 공간에서 이러한 신호들을 제거합니다. 학습된 방향은 훈련 없이도 대조적 디코딩 전략을 통해 환각을 억제하고 일반적인 능력을 유지할 수 있도록 합니다. 세 개의 LVLM 모델에 대해 총 여덟 가지의 평가 지표(환각 및 유용성)를 사용하여 실험한 결과, AOD는 강력한 기준 모델보다 일관되게 우수한 성능을 보였습니다. AOD는 POPE 정확도를 평균 6% 이상 향상시키고, AMBER 점수를 6% 증가시켰으며, MMMU와 같은 유틸리티 작업에서도 뛰어난 성능을 유지했습니다. 추가 분석 결과, AOD는 데이터셋에 국한된 현상이 아닌 일반적인 환각 관련 편향을 포착하며, 다양한 데이터셋에서 강력한 성능을 보이는 것을 확인했습니다. 본 연구의 소스 코드 및 데이터셋은 https://github.com/Hunter-Wrynn/AOD 에서 확인할 수 있습니다.

Original Abstract

Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination-related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training-free dual-forward-pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6\% on average, boosts AMBER by 6\%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination-related biases rather than dataset-specific artifacts. Our source code and datasets are available at https://github.com/Hunter-Wrynn/AOD.

1 Citations
0 Influential
34.431471805599 Altmetric
173.2 Score
Original PDF
3

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!