다중 모드 공간 추론을 위한 시각적 신뢰도 감사
Visual Credit Audit for Multimodal Spatial Reasoning
기존의 폐쇄형 예/아니오(yes/no) 형태의 공간 추론 벤치마크는 이미지 정보가 거의 제공되지 않는 경우에도 올바른 답변에 대해 보상을 제공할 수 있습니다. 본 연구에서는 고정된 선택지 인터페이스 환경에서, 시각적 신뢰도 감사(Visual Credit Audit, VCA)를 통해 두 가지 요소를 분리적으로 평가합니다. 첫째, 벤치마크 이미지의 존재 여부가 모델이 내린 결정에 얼마나 더 많은 정보를 제공하는지를 평가하며, 둘째, 모델이 특정 관계에 대한 시각적 증거에 어떻게 반응하는지를 분석합니다. 첫 번째 평가는 학습 데이터 및 정답 레이블 없이 수행되며, 정답 변경을 요구하지 않습니다. 레이블을 활용하면 '관계 의존성 가중 정확도(dependence-credited correctness, D-CC)'를 얻을 수 있으며, 이는 올바른 항목에서 동일한 제어 그룹과의 비교를 통해 긍정적인 효과를 나타냅니다. 예측 일치성을 확장하여 오답에 대한 감사도 가능합니다. 네 가지 개방형 다중 모드 대규모 언어 모델(MLLM)과 두 개의 공간 추론 벤치마크를 대상으로 분석한 결과, 12.73~26.25%의 결정이 올바르지만 관계 의존성 가중 정확도로 인정되지 않았습니다. 동일 데이터셋 내에서 이미지 순서를 변경하면 D-CC가 21.25~47.80만큼 감소하며, 모든 쌍별 95% 신뢰 구간이 0보다 높습니다. 고정된 픽셀 관계 대조 및 3x3 형태의 증거 소스 요인 분석을 통해, 왜 제어 그룹만으로는 관계 반응을 식별할 수 없는지를 설명합니다. 제어 그룹에 의해 올바르게 판단되었지만 신뢰도가 부여되지 않은 경우, 관계 반전(relation reversal)에 대한 반응은 81.57~100.00%로 높았으며, 전체적으로 32.11%의 응답이 변경되었습니다. 또한, 108개의 기하학적 호환 편집을 통해 독립적인 감사를 수행하여 자연 이미지와의 연관성을 검증했습니다. VCA는 벤치마크 성공을 정확성, 추가적인 이미지 지원 및 관계 일관성 반응이라는 세 가지 요소로 분해합니다.
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.