2608.04575v1 Aug 05, 2026 cs.CV

PhysMind: 비디오를 기반으로 실행 가능한 가상 세계를 구축하여 학습 없이 물리적 추론을 수행하는 방법

PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

Youquan He
Youquan He
Citations: 3
h-index: 1
Mingyi Deng
Mingyi Deng
Citations: 21
h-index: 3
Shenxiang Zeng
Shenxiang Zeng
Citations: 1
h-index: 1
Haoyang Zhao
Haoyang Zhao
Citations: 16
h-index: 2
Chen Yang
Chen Yang
Citations: 90
h-index: 4
Zhouyuan Xu
Zhouyuan Xu
Citations: 4
h-index: 2
Haoyu Li
Haoyu Li
Citations: 9
h-index: 1
Jia Fan
Jia Fan
Citations: 67
h-index: 1
Chen Wang
Chen Wang
Citations: 41
h-index: 3

비디오로부터 신뢰성 있는 물리적 추론을 위해서는 객체의 움직임, 상호 작용, 그리고 외부 개입에 대한 반응을 이해해야 합니다. 기존의 시각-언어 모델(VLM)은 이러한 역학 관계를 해석하고 미래 및 가상 상황에 대해 안정적으로 추론하는 데 어려움을 겪는 경우가 많습니다. 본 논문에서는 PhysMind라는 학습이 필요 없는 에이전트 기반 프레임워크를 소개합니다. PhysMind는 각 비디오에 대해 재사용 가능하며 질문에 독립적인 실행 가능한 가상 세계를 구축합니다. PhysMind는 객체 분할, 메시 복원, 그리고 6자유도(6D) 자세 추적을 통해 시간적으로 일관된 동적 장면을 복구한 다음, 시간 기반 시뮬레이터를 사용하지 않고 분석적인 연속 시간 역학과 잠재적인 물리적 매개변수를 추정합니다. 질문이 주어지면 PhysMind는 가상 세계를 검사하고, 확장하거나 수정하며, 결과적으로 발생하는 궤적과 상호 작용을 바탕으로 답변을 생성합니다. 동일한 VLM을 사용한 직접적인 연쇄적 사고(CoT) 추론에 비해, PhysMind는 CLEVRER 데이터셋에서 38.23점, Physion++ 데이터셋에서 8.08점을 향상된 정확도를 보입니다. 가상 질문에 대한 답변에서는 평가된 가장 강력한 VLM인 GPT-5.5를 19.25점 이상 능가합니다.

Original Abstract

Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic scene through object segmentation, mesh reconstruction, and 6D pose tracking, then fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator. Given a question, it inspects, continues, or edits the world and answers from the resulting trajectories and interactions. Relative to direct chain-of-thought (CoT) reasoning with the same VLM, PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++. On counterfactual questions, it exceeds the strongest evaluated VLM baseline, GPT-5.5, by 19.25 points.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!