2606.16494v1 Jun 15, 2026 cs.CL

결국에는 사라진다: 다중 모드 검색 증강 질의응답 시스템에서의 선호 편향

Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering

Zhen Wang
Zhen Wang
Citations: 11
h-index: 3
Jieyuan Liu
Jieyuan Liu
Citations: 9
h-index: 2
Jianyang Gu
Jianyang Gu
Citations: 84
h-index: 4
Shijie Chen
Shijie Chen
Citations: 22
h-index: 2
Jefferson Chen
Jefferson Chen
Citations: 8
h-index: 2

지식 기반 시각 질의 응답 (KB-VQA)은 시각-언어 시스템이 파라미터 지식을 초월하는 질문에 답할 수 있도록, 위키피디아 규모의 지식 베이스에서 검색된 문서를 활용하여 답변을 생성합니다. 순수 텍스트 기반의 장문 맥락 LLM에서 검색된 맥락을 사용할 때, Liu et al. (2024)이 밝힌 U자형 '중간 부분 손실' 효과가 나타납니다. 즉, 맥락의 시작과 끝부분 정보는 활용되지만 중간 부분이 잊혀집니다. 이러한 현상이 실제 다중 모드 KB-VQA 시스템에서도 발생하는지 여부는 아직 명확하지 않습니다. 이 연구에서는 다중 모드 KB-VQA 시스템에서 Reader 측 위치 의존성을 처음으로 체계적으로 분석하기 위해, '골드 포지션 프로토콜'을 설계했습니다. 이 프로토콜은 질문 내에서 오직 '골든 패시지(정답 문맥)'에 해당하는 프롬프트 슬롯만 변경합니다. 우리는 세 가지 오픈 소스 7B/8B VLM Reader와 두 개의 KB-VQA 벤치마크를 사용하여 k 값이 최대 20까지의 실험을 진행했습니다. 결과는 기존 U자형 패턴에서 '시작 부분 선호' 현상으로 바뀌었으며, 모든 Reader-벤치마크 조합에서 '골든 패시지'가 시작 부분에 위치할 때 마지막 부분보다 16~26점 더 높은 성능을 보였습니다. 우리는 이 효과를 '결국에는 사라진다 (Lost at the End)'라고 명명했습니다. 세 가지 실험을 통해 원인을 분석한 결과, 텍스트만으로 구성된 제어 그룹에서 다중 모드 설정이 기존 텍스트 기반의 선호 현상을 최대 2.2배에서 4.5배까지 증폭시키는 것을 확인했습니다. 또한 이미지 위치 변경 및 방해 요소 재정렬 실험을 통해 원인이 Instruction-tuned Reader의 프롬프트 슬롯 0에 있음을 밝혀냈습니다. frozen된 Reader 상태에서 세 가지 검색 측면 개선 방법 (MMR, 오라클 재순위화, 순위 기반 재정렬)을 적용했지만, 성능 차이를 줄이는 데 실패했습니다. 우리의 연구 결과는 recall@k가 실제 KB-VQA 시스템 평가에 적합한 지표가 아니며, 성능 향상을 위해서는 Reader 측면의 개입이 필요하다는 것을 시사합니다. 우리는 본 연구에서 개발한 프로토콜을 이러한 개선 노력을 평가하는 데 활용할 수 있는 도구로 공개합니다.

Original Abstract

Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by conditioning a reader on passages retrieved from a Wikipedia-scale knowledge base. In pure-text long-context LLMs, retrieved-context use follows the U-shaped "lost-in-the-middle" effect of Liu et al. (2024): information at the start and end of context is used, the middle is lost. Whether this transfers to deployed multimodal KB-VQA is open. To close this gap, we design the first controlled probe of reader-side position dependence in multimodal KB-VQA: a gold-position protocol in which only the gold passage's prompt slot varies within question. We run it on three open-source 7B/8B VLM readers and two KB-VQA benchmarks at k up to 20. The shape flips from U to primacy: gold-at-first beats gold-at-last by 16 to 26 points on every reader-by-benchmark cell, an effect we call "Lost at the End". Three targeted ablations narrow the cause: a text-only control shows the multimodal setting amplifies an already-present text-mode primacy 2.2 to 4.5 times, and image-position and distractor-shuffle ablations together pin the locus to prompt slot 0 of the instruction-tuned reader. On a frozen reader, three retrieval-side fixes (MMR, oracle reranking, rank-based reordering) all leave the gap intact (no separable improvement). Our findings indicate that recall@k is the wrong metric for deployed KB-VQA and that closing the gap requires reader-side intervention; we release our protocol as a controlled instrument for evaluating such interventions.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!