물리적 추론을 위한 인과적 가이드: 시각-언어 모델(VLM)의 인과 기반 물리 세계 이해를 위한 벤치마크
Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs
물리 세계에 대한 이해와 추론은 지능적인 행동의 기본이지만, 최첨단 시각-언어 모델(VLM)은 여전히 인과적 물리 추론에 실패하는 경우가 많으며, 종종 그럴듯하지만 틀린 답변을 생성합니다. 이러한 격차를 해결하기 위해, 우리는 3,000개 이상의 신중하게 선별된 비디오 및 이미지 기반 질문으로 구성된 벤치마크인 CausalPhys를 소개합니다. 이 벤치마크는 인지, 예측, 개입, 목표 지향의 네 가지 영역을 포괄합니다. 각 질문은 객체-속성-사건 간의 의존성을 나타내는 전문가가 주석을 달아 생성한 인과 그래프와 함께 제공되어, 인과적 이해에 대한 해석 가능하고 세분화된 평가를 가능하게 합니다. 이를 바탕으로, 우리는 모델의 추론 과정이 올바른 인과 관계와 얼마나 일치하는지를 정량적으로 측정하는 인과 그래프 기반 지표를 제안합니다. 이 지표는 단순히 답변 정확도만을 측정하는 것이 아니라, VLM의 인과적 추론 실패에 대한 체계적인 진단을 가능하게 합니다. 이 지표를 사용하여 주요 VLM에 대한 종합적인 분석을 수행한 결과, 객체 간의 인과적 의존성을 파악하는 데 있어 일관된 격차가 존재하며, 이는 인과 관계를 고려한 학습의 필요성을 강조합니다. 이러한 한계를 극복하기 위해, 우리는 VLM의 추론 과정을 명시적으로 인과 구조와 일치시키는 Causal Rationale-informed Fine-Tuning (CRFT)이라는 방법을 제안합니다. 광범위한 실험 결과, CRFT는 다양한 모델 아키텍처에서 추론 정확도와 해석 가능성을 크게 향상시키는 것으로 나타났습니다. CausalPhys는 데이터셋 큐레이션, 인과적 평가 및 인과 관계를 고려한 학습을 통합하여, 현대 VLM이 인과적으로 근거한 물리적 추론으로 발전하는 데 중요한 기반을 제공합니다.
Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) still fail at causal physical reasoning, often producing plausible but incorrect answers. To address this gap, we introduce CausalPhys, a benchmark of over 3,000 carefully curated video- and image-based questions spanning four domains: Perception, Anticipation, Intervention, and Goal Orientation. Each question is paired with an expert-annotated causal graph capturing object-attribute-event dependencies, enabling interpretable and fine-grained evaluation of causal understanding. Building on this, we formulate a causal-graph-grounded metric that quantitatively measures how well a model's chain-of-thought reasoning aligns with the correct causal relations, moving beyond answer-only accuracy and enabling systematic diagnosis of VLMs' causal reasoning failures. Using this metric, we conduct a comprehensive analysis of leading VLMs, revealing systematic gaps in capturing causal dependencies and underscoring the need for causality-aware learning. To address these limitations, we further propose Causal Rationale-informed Fine-Tuning (CRFT), which explicitly aligns VLM reasoning with causal structures. Extensive experiments demonstrate that CRFT substantially enhances both reasoning accuracy and interpretability across multiple model backbones. By unifying dataset curation, causal evaluation, and causality-informed learning, CausalPhys establishes a strong foundation for advancing modern VLMs toward causally grounded physical reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.