비디오 기반 모델이 직관적인 물리학을 이해하는가? 계층별 탐색 분석
Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis
본 연구에서는 사전 학습된 비디오 기반 모델들이 고정된 표현 내에 직관적인 물리학 정보를 얼마나 포함하고 있는지, 그리고 이러한 정보가 모델 종류, 계층, 탐색 방식에 따라 어떻게 달라지는지를 조사합니다. IntPhys2 및 Minimal Video Pairs (MVP) 데이터셋을 사용하여 고정된 특징 탐색(frozen-feature probing) 방법을 통해 예측 기반의 결합 임베딩 모델(V-JEPA), 마스크 복원 모델(VideoMAE), 그리고 확산 기반 비디오 생성 모델(LTX-Video)을 비교했습니다. V-JEPA는 전반적으로 가장 우수한 성능을 보였으며, 특히 시간적 동역학을 모델링하는 탐색 방식에서 뛰어난 결과를 나타냈습니다. VideoMAE는 경쟁력 있는 성능을 유지했으며, LTX-Video는 상대적으로 약하지만 의미있는 신호를 회복했습니다. 계층별 분석 결과, 물리학과 관련된 정보는 초기 계층에서 가장 취약하며, 중간~후반 계층에서 가장 쉽게 접근할 수 있음을 확인했습니다. 시간 제어 실험에서는 프레임 순서를 변경하면 성능이 크게 저하되는 것을 보여주었으며, 특히 MVP 데이터셋에서 두드러졌습니다. 이러한 결과를 종합적으로 고려할 때, 사전 학습된 비디오 표현 내에 직관적인 물리학 지식이 안정적으로 나타나지만, 접근성은 사전 훈련 방식, 표현의 깊이 및 출력 메커니즘에 크게 의존하는 것으로 보입니다.
We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the strongest overall results across benchmarks, especially with probes that model temporal dynamics, while VideoMAE remains competitive and LTX-Video recovers weaker but non-trivial signal. Layerwise analyses show that physics-relevant information is weakest in early layers and becomes most accessible at intermediate-to-late depth, and temporal controls show that disrupting frame order substantially reduces performance, especially on MVP. Together, these results suggest that intuitive-physics knowledge emerges reliably in pretrained video representations, but its accessibility depends strongly on pretraining paradigm, representational depth, and readout mechanism.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.