PlayWorld: 자율적인 놀이를 통해 학습하는 로봇 세계 모델
PlayWorld: Learning Robot World Models from Autonomous Play
행동에 조건화된 비디오 모델은 데이터로부터 직접 개선될 수 있는 범용 로봇 시뮬레이터를 구축하는 유망한 방법을 제공합니다. 그러나 대규모 로봇 데이터셋으로 학습했음에도 불구하고, 현재 최고 성능의 비디오 모델은 여전히 로봇 조작에 중요한 물리적으로 일관된 로봇-객체 상호작용을 예측하는 데 어려움을 겪습니다. 이러한 격차를 해소하기 위해, 우리는 상호작용 경험으로부터 고품질 비디오 세계 시뮬레이터를 학습시키는 간단하고 확장 가능하며 완전 자율적인 파이프라인인 PlayWorld를 제시합니다. 기존 접근 방식이 성공 편향된 인간 시연에 의존하는 것과는 달리, PlayWorld는 인간의 개입 없이 로봇의 자체 학습만을 통해 학습할 수 있는 최초의 시스템입니다. 이를 통해 복잡하고 드물게 발생하는 물리적 상호작용을 자연스럽게 수집하고, 현실적인 객체 역학을 모델링하는 데 필수적입니다. 다양한 조작 작업에 대한 실험 결과, PlayWorld는 인간이 수집한 데이터로 학습된 세계 모델이 포착하지 못하는 접촉이 많은 상호작용에 대해 고품질의 물리적으로 일관된 예측을 생성합니다. 또한, PlayWorld는 세밀한 실패 예측 및 정책 평가를 가능하게 하며, 인간이 수집한 데이터보다 최대 40%의 성능 향상을 보입니다. 마지막으로, PlayWorld가 세계 모델 내에서 강화 학습을 가능하게 하여, 실제 환경에 적용했을 때 성공률을 65% 향상시키는 정책 성능 향상을 보여줍니다.
Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still struggle to predict physically consistent robot-object interactions that are crucial in robotic manipulation. To close this gap, we present PlayWorld, a simple, scalable, and fully autonomous pipeline for training high-fidelity video world simulators from interaction experience. In contrast to prior approaches that rely on success-biased human demonstrations, PlayWorld is the first system capable of learning entirely from unsupervised robot self-play, enabling naturally scalable data collection while capturing complex, long-tailed physical interactions essential for modeling realistic object dynamics. Experiments across diverse manipulation tasks show that PlayWorld generates high-quality, physically consistent predictions for contact-rich interactions that are not captured by world models trained on human-collected data. We further demonstrate the versatility of PlayWorld in enabling fine-grained failure prediction and policy evaluation, with up to 40% improvements over human-collected data. Finally, we demonstrate how PlayWorld enables reinforcement learning in the world model, improving policy performance by 65% in success rates when deployed in the real world.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.