2608.04244v1 Aug 04, 2026 cs.CV

SIGNPOST-Bench: 다중 모드 대규모 언어 모델에서의 텍스트-이미지 충돌 해결 성능 평가

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Ling Dai
Ling Dai
Citations: 19
h-index: 2
Junting Zhou
Junting Zhou
Citations: 439
h-index: 7
Minghao Liu
Minghao Liu
Citations: 293
h-index: 6
Sirun Li
Sirun Li
Citations: 187
h-index: 1
Haoxin Lyu
Haoxin Lyu
Citations: 8
h-index: 2
Yong Li
Yong Li
Citations: 0
h-index: 0
Fan Zhang
Fan Zhang
Citations: 72
h-index: 5

다중 모드 대규모 언어 모델(MLLM)은 시각적 단서와 텍스트 정보를 결합하여 실제 환경에서 정확한 예측을 수행하지만, 기존의 평가 지표들은 이러한 정보들이 충돌할 때 모델이 어떻게 판단하는지를 제대로 보여주지 못합니다. 본 연구에서는 텍스트-이미지 충돌 해결 능력을 평가하기 위한 제어된 반사실적 평가 도구인 SIGNPOST-Bench를 소개합니다. 각 원본 이미지는 Original, Blank(빈칸), Similar(유사), Random(랜덤), 그리고 Adversarial(적대적)의 다섯 가지 변형으로 구성됩니다. 합성된, 특정 영역에 국한된 텍스트 변화는 시각적 내용을 보존하도록 설계되었으며, 이를 통해 위치 추정 성능의 변화와 충돌하는 텍스트가 도입되었을 때 발생하는 지리적 목표 방향 변화를 측정합니다. SIGNPOST-Bench는 네 개의 데이터 세트에서 추출된 5,111개의 반사실적 그룹과 25,555개의 이미지 변형으로 구성되어 있습니다. 본 연구에서는 총 20개의 MLLM을 7개 제공업체로부터 평가했습니다. 원본 이미지와 비교했을 때, 적대적 변형은 중앙값 기준 위치 추정 오류를 282km에서 1,347km로 증가시켰으며, 이는 4.8배의 증가입니다. 평가 대상 모델 중에서, 지리 정보가 포함된 적대적 샘플에 대해 6.5%에서 20.1%의 예측 결과가 주입된 목표 위치로부터 50km 이내에 있었습니다. 또한, 모든 평가 모델에서 Blank 변형 대비 Adversarial 변형으로 인해 목표 거리 감소량이 양호했습니다. 일관성 있는, 관련 없는, 그리고 충돌하는 텍스트 대체는 모델 예측에 서로 다른 영향을 미치며, 원본 이미지에서의 위치 추정 성능은 충돌하는 텍스트에 대한 모델의 강건성을 완전히 예측하지 못합니다. 이러한 결과는 시각적 지리 정보가 장면 텍스트 해결 능력을 진단하는 데 유용하며, MLLM이 충돌하는 다중 모드 정보를 어떻게 처리하는지를 평가하기 위한 제어된 프레임워크를 제공합니다.

Original Abstract

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!