HIVE: 시각 언어 모델에서 환각 이후의 추론 이해
HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models
시각 언어 모델(VLM)에서의 환각은 일반적으로 의미 오류로 간주되지만, 종종 불완전하거나 모호한 시각적 증거에서 비롯됩니다. 기존 연구는 주로 생성 단계에서 환각을 감지하거나 억제하는 데 초점을 맞추고 있으며, 이후의 추론 단계는 상대적으로 탐구되지 않았습니다. 본 논문에서는 환각된 의미가 모델의 추론 맥락에 진입하여 후속 예측에 영향을 미치는 단계인 '환각 이후 추론(Post Hallucination Reasoning, PHR)'을 연구합니다. 체계적인 PHR 조사를 위해, 신뢰할 수 있는 캡션과 환각된 캡션을 비교할 수 있는 평가 인프라인 HIVE (Hallucination Inference and Verification Engine)를 소개합니다. 9개의 작업과 9개의 모델에 대한 분석 결과, 모달리티 의존적인 패턴이 관찰되었습니다. 환각된 캡션은 시각 언어 작업의 정확도를 향상시키는 경향이 있지만, 텍스트 기반 작업에서는 제한적이거나 불안정한 영향을 보이는 경우가 많습니다. 추가 분석을 통해 환각된 단서는 의미 범위를 확장하고 추론 역학을 재구성하는 동시에 안정적인 추론을 유지함을 확인했습니다. 이러한 결과는 환각된 의미가 모델의 추론 맥락에 진입하면 후속 추론에 영향을 미칠 수 있음을 시사합니다. 이 '환각 이후' 단계에 대한 이해는 다중 모달 추론 시스템의 신뢰성과 해석 가능성을 향상시키는 데 중요합니다. 코드 및 관련 자료는 다음 링크에서 공개적으로 이용할 수 있습니다: https://github.com/hefengcs/HIVE.
Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence. Prior work mainly focuses on detecting or suppressing hallucinations at generation time, leaving the subsequent reasoning stage largely unexplored. In this work, we study Post Hallucination Reasoning (PHR), the stage in which hallucinated semantics enter the model's inference context and influence downstream predictions. To systematically investigate PHR, we introduce HIVE, Hallucination Inference and Verification Engine, an evaluation infrastructure that enables controlled comparisons between faithful and hallucinated captions. Across nine tasks and nine models, we observe structured modality dependent patterns: hallucinated captions often improve accuracy on vision language tasks, while text only tasks exhibit limited or unstable effects. Further analyses show that hallucinated cues broaden semantic coverage and reshape reasoning dynamics while preserving stable inference. These findings highlight that hallucinated semantics may influence downstream reasoning once they enter the model's inference context. Understanding this post hallucination stage is important for improving the reliability and interpretability of multimodal reasoning systems. Code is publicly available at https://github.com/hefengcs/HIVE.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.