PhiZero: 물리적 언어를 기반으로 구축된 세계 모델
PhiZero: A World Model Built Around Physical Language
본 논문에서는 물리적 언어를 기반으로 구축된 세계 모델인 PhiZero를 소개합니다. 물리적 언어는 세상의 상태 변화를 간결하게 표현하는 이산적인 형태입니다. 기존의 물리적 세계 모델은 일반적으로 픽셀 공간에서 미래 영상을 직접 예측하는데, 이는 근본적인 세계 역학을 고차원 시각 예측기에 암묵적으로 내재시켜 버립니다. 인간이 시각 경험으로부터 예측 구조를 추상화하고 이를 자연어로 조직하여 명시적인 추론을 수행하는 능력에 영감을 받아, 우리는 자체 감독 학습을 통해 실제 영상에서 물리적 언어를 학습하고, 이를 사용하여 물리 세계가 어떻게 변화하는지에 대해 명시적으로 추론합니다. 이에 따라 PhiZero는 '추론-렌더링' 패러다임을 채택합니다: 먼저 미래의 세계 변화를 물리적 언어 시퀀스로 추론한 다음, 추론된 변화를 영상으로 렌더링합니다. 생성 및 이해 성능 평가 지표에 대한 광범위한 실험을 통해 PhiZero가 물리적으로 일관성 있는 세계 변화를 모델링하는 능력을 검증했습니다. 또한, PhiZero는 현실적이고 상호작용적인 세계 모델링, 세분화된 동작 기반 시뮬레이션, 그리고 제로샷 모션 전송에 대한 잠재력을 보여줍니다.
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.