PhysMani: 물리학 기반 3차원 세계 모델을 이용한 동적 객체 조작
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
구조화되지 않은 3차원 환경에서 빠르게 움직이는 대상을 조작하는 것은 로봇 인공지능에게 여전히 어려운 과제입니다. 기존의 시각-언어-행동 모델과 세계 모델은 정확한 3차원 기하 정보와 물리적으로 의미 있는 예측 능력이 부족합니다. 본 논문에서는 물리학 기반 3차원 가우시안 세계 모델과 미래를 고려한 행동 정책 모델을 결합한 PhysMani 프레임워크를 제안합니다. 제안하는 세계 모델은 온라인 최적화를 통해 발산이 없는 가우시안 속도장을 학습하여 빠르고 물리적으로 정확한 미래 동역학 예측을 수행합니다. 또한, 정책 모델은 학습 가능한 토큰 기반 크로스 어텐션 모듈을 통해 예측된 3차원 장면의 미래 동역학 정보를 통합합니다. 우리는 16개의 작업으로 구성된 동적 조작 벤치마크인 PhysMani-Bench를 소개하고, 시뮬레이션 및 실제 로봇 실험에서 강력한 기존 모델보다 우수한 성공률을 달성했음을 보여줍니다.
Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models struggle with accurate 3D geometry and physically meaningful forecasting. We propose PhysMani, a framework that couples a physics-principled 3D Gaussian world model with a future-aware action policy model. The world model learns a divergence-free Gaussian velocity field via online optimization for fast and physically grounded future dynamics prediction. The policy model integrates the predicted 3D scene future dynamics through a learnable token based cross-attention module. We introduce PhysMani-Bench, a dynamic manipulation benchmark with 16 tasks, and demonstrate a superior success rate over strong baselines in both simulation and real-world robot experiments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.