MindEdit-Bench: 야생 환경 사진을 활용한 시각-언어 모델의 객체 수준 반사실적 공간 추론 성능 평가
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
시각-언어 모델(VLM)에 대한 기존의 성능 평가는 주로 관찰 기반의 공간 추론 능력을 측정하며, 모델은 입력 이미지에 이미 존재하는 관계를 설명합니다. 기존의 '만약 ~라면' 유형의 작업은 주로 장면 자체는 고정하고 관찰자의 위치를 변경하는 방식입니다. 본 연구에서는 VLM이 객체를 가상으로 이동하거나 회전시켰을 때 발생하는 결과를 예측할 수 있는지 탐구합니다. 우리는 새로 촬영된 실내 장면을 활용하여 3D 장면 그래프 추출 파이프라인을 통해 구축된, 6가지 공간 추론 작업을 포함하는 성능 평가 도구인 MindEdit-Bench를 소개합니다. 이 중 4가지 작업은 관찰 가능한 구조에 대한 인식 및 원근 변환 능력을 평가하며, 2가지 새로운 작업(L4: 공간 편집, L5: 다중 시점 가시성 편집)은 객체 수준의 반사실적 추론 능력을 평가합니다. 이때 정답이 입력 이미지 어디에도 존재하지 않습니다. 각 질문에는 8개에서 24개의 구조화된 답변 선택지가 제공되어, 공간 추론 및 오류 발생 원인을 세밀하게 분석할 수 있습니다. 이 성능 평가 도구는 공개 데이터셋에서 추출되지 않은 120개의 비공개 실내 장면을 포함하여, 공개 데이터 기반 사전 학습으로 인한 편향 위험을 줄였습니다. 1,003개의 인간 검증된 질문에 대해 15개의 VLM을 평가한 결과, 작업별 평균 VLM 정확도는 8%에서 31%에 불과했습니다. 반면, 인간의 다수 투표 정확도는 81%에서 97%였습니다. 전체적으로 인간과 최적 VLM 간의 성능 격차는 53pp (percentage points)였으며, 모든 작업에서 최소 39pp 이상의 격차가 발생했습니다. 구조화된 답변 공간 분석 결과, 카메라-심도 축 추론 능력 부족 및 어려운 가시성 편집 문제에 대한 오류 처리 방식과 같은 비균일적인 실패 패턴이 나타났습니다.
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-the-wild 3D scene-graph extraction pipeline. Four tasks probe perception and perspective transformation over observed structure; two new tasks, L4 (spatial editing) and L5 (cross-view visibility editing), probe object-level counterfactual reasoning, where correct answers are absent from all input images. Each question provides 8-24 structured answer choices, enabling answer-letter-level diagnosis of spatial and fallback errors. The benchmark covers 120 private indoor scenes not drawn from public datasets, reducing public-data pretraining-overlap risk. Across 15 VLMs on 1,003 human-verified questions, task-wise mean VLM accuracy is only 8%-31%, versus 81%-97% human majority-vote accuracy. The pooled human--best-VLM gap is 53 pp, with at least 39 pp on every task. The structured answer space further reveals non-uniform failures, including weaker camera-depth-axis inference and fallback behavior on difficult visibility-editing cases.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.