2607.27069v1 Jul 29, 2026 cs.CV

다중 모드 공간 추론을 위한 시각적 신뢰도 감사

Visual Credit Audit for Multimodal Spatial Reasoning

Xueqi Cheng
Xueqi Cheng
Citations: 1,916
h-index: 22
Huawei Shen
Huawei Shen
Citations: 422
h-index: 13
Nan Wei
Nan Wei
Citations: 3,741
h-index: 5
Qiang Qiu
Qiang Qiu
Citations: 115
h-index: 4
Feixiang Liu
Feixiang Liu
Citations: 19
h-index: 2
Lanbo Sun
Lanbo Sun
Citations: 0
h-index: 0

기존의 폐쇄형 예/아니오(yes/no) 형태의 공간 추론 벤치마크는 이미지 정보가 거의 제공되지 않는 경우에도 올바른 답변에 대해 보상을 제공할 수 있습니다. 본 연구에서는 고정된 선택지 인터페이스 환경에서, 시각적 신뢰도 감사(Visual Credit Audit, VCA)를 통해 두 가지 요소를 분리적으로 평가합니다. 첫째, 벤치마크 이미지의 존재 여부가 모델이 내린 결정에 얼마나 더 많은 정보를 제공하는지를 평가하며, 둘째, 모델이 특정 관계에 대한 시각적 증거에 어떻게 반응하는지를 분석합니다. 첫 번째 평가는 학습 데이터 및 정답 레이블 없이 수행되며, 정답 변경을 요구하지 않습니다. 레이블을 활용하면 '관계 의존성 가중 정확도(dependence-credited correctness, D-CC)'를 얻을 수 있으며, 이는 올바른 항목에서 동일한 제어 그룹과의 비교를 통해 긍정적인 효과를 나타냅니다. 예측 일치성을 확장하여 오답에 대한 감사도 가능합니다. 네 가지 개방형 다중 모드 대규모 언어 모델(MLLM)과 두 개의 공간 추론 벤치마크를 대상으로 분석한 결과, 12.73~26.25%의 결정이 올바르지만 관계 의존성 가중 정확도로 인정되지 않았습니다. 동일 데이터셋 내에서 이미지 순서를 변경하면 D-CC가 21.25~47.80만큼 감소하며, 모든 쌍별 95% 신뢰 구간이 0보다 높습니다. 고정된 픽셀 관계 대조 및 3x3 형태의 증거 소스 요인 분석을 통해, 왜 제어 그룹만으로는 관계 반응을 식별할 수 없는지를 설명합니다. 제어 그룹에 의해 올바르게 판단되었지만 신뢰도가 부여되지 않은 경우, 관계 반전(relation reversal)에 대한 반응은 81.57~100.00%로 높았으며, 전체적으로 32.11%의 응답이 변경되었습니다. 또한, 108개의 기하학적 호환 편집을 통해 독립적인 감사를 수행하여 자연 이미지와의 연관성을 검증했습니다. VCA는 벤치마크 성공을 정확성, 추가적인 이미지 지원 및 관계 일관성 반응이라는 세 가지 요소로 분해합니다.

Original Abstract

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!