2608.03782v1 Aug 04, 2026 cs.AI

KnowHal: 지식 기반의 종합적인 다중 모드 환각 평가 벤치마크

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

Yu Huang
Yu Huang
Citations: 30
h-index: 4
Huining Li
Huining Li
Citations: 1
h-index: 1
Qian Li
Qian Li
Citations: 9
h-index: 1
Yuntao Du
Yuntao Du
Citations: 469
h-index: 5
Kailin Jiang
Kailin Jiang
Citations: 90
h-index: 6
Ruihan Li
Ruihan Li
Citations: 0
h-index: 0
Jiyang Tan
Jiyang Tan
Citations: 0
h-index: 0
Hengyang Lu
Hengyang Lu
Citations: 11
h-index: 1

환각 현상은 신뢰할 수 있는 다중 모드 대규모 언어 모델(MLLM)을 개발하는 데 있어 중요한 과제입니다. 기존 벤치마크는 주로 개체, 속성 및 관계 환각에 초점을 맞추고 있지만, 지식 관련 오류는 종종 별도로 조사되며, 다양한 환각 측면 전반에 걸친 통합적인 평가 프레임워크가 부족합니다. 이러한 문제를 해결하기 위해, 우리는 다중 모드 환각 평가에서 개체, 속성, 관계 및 지식의 네 가지 차원에 걸쳐 명시적으로 지식 환각을 포함하는 벤치마크인 **KnowHal**을 제안합니다. KnowHal은 공유 이미지와 개체에 대한 쌍으로 구성된 긍정적 및 부정적인 질문을 생성하여, 인식 오류, 지식 관련 오류 및 허위 전제 수용 간의 통제된 비교를 가능하게 합니다. 이 벤치마크는 LLM 지원, CLIP 기반 필터링 및 인간 검증을 결합한 반자동 파이프라인을 통해 구성된 10개 도메인 및 50개 범주의 1,800개의 샘플로 구성되어 있습니다. 우리는 KnowHal에서 14개의 대표적인 MLLM을 평가하고 광범위한 분석을 수행했습니다. 결과는 지식 차원이 거의 모든 모델에 대해 가장 큰 어려움을 나타내는 것을 보여주었으며, 대부분의 모델이 부정적인 질문에 대해 상당한 성능 저하를 보이는 것으로 나타나 허위 전제에 대한 제한된 견고성을 드러냈습니다. KnowHal은 쌍으로 구성된 질문 설계를 통해 네 가지 환각 차원을 통합함으로써 기존 평가 프레임워크의 중요한 격차를 해소하고 MLLM에서의 환각 현상에 대한 보다 종합적인 평가를 가능하게 합니다.

Original Abstract

Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus on entity, attribute, and relation hallucinations, knowledge-related failures are often investigated separately, lacking a unified evaluation framework across different hallucination dimensions. To overcome this, we propose \textbf{KnowHal}, a benchmark that explicitly incorporates knowledge hallucination into multimodal hallucination evaluation spanning four dimensions: entity, attribute, relation, and knowledge. KnowHal constructs paired positive and negative questions over shared images and entities, enabling controlled comparisons among perceptual errors, knowledge-related errors, and false-premise acceptance. The benchmark contains 1,800 samples across 10 domains and 50 categories, constructed through a semi-automated pipeline combining LLM assistance, CLIP-based filtering, and human verification. We evaluate 14 representative MLLMs on KnowHal and conduct extensive analyses. Results show that the knowledge dimension consistently presents the greatest challenge for nearly all evaluated models, while most models exhibit substantial performance degradation on negative questions, revealing limited robustness to false premises. By unifying four hallucination dimensions with paired question design, KnowHal addresses an important gap in existing evaluation frameworks and enables a more comprehensive assessment of hallucinations in MLLMs.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!