2607.21155v1 Jul 23, 2026 cs.CV

CRAG-MM-Diagnostics: 지식 집약형 시각 질의응답(VQA) 시스템의 단계별 분석을 위한 진단 도구

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Siva Reddy
Siva Reddy
Citations: 10,990
h-index: 46
Hanseok Oh
Hanseok Oh
Citations: 184
h-index: 6
Parishad BehnamGhader
Parishad BehnamGhader
Citations: 852
h-index: 6
Benno Krojer
Benno Krojer
Citations: 21
h-index: 3
Hyunji Lee
Hyunji Lee
Citations: 223
h-index: 9
Paul Liang
Paul Liang
Citations: 0
h-index: 0
Verna Dankers
Verna Dankers
University of Edinburgh
Citations: 1,028
h-index: 14

지식 집약형 시각 질의응답(KI-VQA) 벤치마크는 외부 정보를 활용하여 이미지에 대한 질문에 답하도록 함으로써, 비전-언어 모델(VLM)을 다중 모드 지식 어시스턴트로서 평가합니다. KI-VQA는 표현 이해, 시각적 참조 연결, 객체 인식, 지식 검색 및 추론 등 여러 하위 문제를 포함하지만, 기존 벤치마크는 일반적으로 최종 작업의 정확도만을 보고하여 어떤 부분에서 오류가 발생하는지 파악하기 어렵습니다. KI-VQA 전체 프로세스를 분석하기 위해, 우리는 단계별 데이터 주석을 통해 1) 언어 기반 시각적 참조 연결, 2) 객체 식별, 그리고 3) 지식 검색 및 추론을 분리하는 진단 벤치마크인 CRAG-MM-Diagnostics를 소개합니다. 우리는 완전하게 파라미터화된 모델과 검색 증강 VLM을 평가하고, 새로 수집된 메타데이터(예: 대상 ROI, 개체 이름, 시각적 복잡도 점수)를 사용하여 세밀한 분석을 제공합니다. 우리의 결과는 지식 검색 및 추론이 주요 병목 현상임을 보여주지만, KI-VQA 파이프라인의 다른 부분에서도 문제가 있음을 강조합니다. 예를 들어, VLM은 대상 객체 식별에 어려움을 겪거나, 이미지 검색 시스템이 텍스트 단서를 통합하는 데 어려움을 겪을 수 있습니다. 이러한 결과는 현재 KI-VQA 시스템의 근본적인 한계를 드러내고 단계별 평가의 필요성을 강조합니다. 마지막으로, 우리는 이러한 결과를 활용하여 시각적 참조 연결 모듈을 통합하여 이미지 검색 전에 대상을 잘라내는 접지된 양방향 RAG 파이프라인을 제안했습니다. 이 방법은 GPT-5와 Qwen의 정확도를 각각 13.3% 및 8.5% 포인트 향상시켰습니다.

Original Abstract

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise. To analyze the full KI-VQA pipeline, we introduce CRAG-MM-Diagnostics, a diagnostic benchmark with stage-wise data annotations that isolate 1) language-based visual grounding, 2) object identification, and 3) knowledge retrieval and reasoning. We evaluate fully parametric and retrieval-augmented VLMs, providing fine-grained analyses using newly collected metadata, such as target ROIs, entity names, and visual complexity scores. Our results point to knowledge retrieval and reasoning as the primary bottleneck, but also highlight issues in the other parts of the KI-VQA pipeline, such as the fact that VLMs struggle with target object identification or that image retrievers struggle to integrate textual cues. These findings expose fundamental limitations in current KI-VQA systems and motivate stage-aware evaluation. We, lastly, leverage these findings to propose a grounded bimodal RAG pipeline that integrates a visual grounding module to crop targets before image retrieval, boosting GPT-5 and Qwen's respective accuracies by 13.3 and 8.5 percentage points.

0 Citations
0 Influential
23 Altmetric
115.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!