GeniWorld: 시각적 액션을 활용한 로봇 조작을 위한 일반화 가능한 상호 작용형 세계 모델
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
일반적인 로봇 정책은 강력한 기능을 보이지만, 복잡하고 새로운 환경에서의 견고성은 여전히 제한적입니다. 다양한 실제 환경에서 로봇 학습 및 평가를 확장하는 것은 비용이 많이 들고 어려운 과제입니다. 액션 기반 세계 모델은 유망한 대안을 제공하지만, 종종 제한된 액션 제어 능력과 데이터 분포 외부(out-of-distribution) 시나리오에 대한 낮은 일반화 성능 문제를 겪습니다. 이에 따라, 우리는 다양한 새로운 환경에서도 견고하게 일반화되는 로봇용 상호 작용형 세계 모델인 GeniWorld를 제시합니다. 사전 학습된 비디오 생성 모델을 기반으로, URDF 기반 렌더링을 사용하여 수치 액션을 시각적 액션 표현으로 변환하여 공간적으로 제어 가능한 액션을 가능하게 합니다. 당사 모델은 로봇의 운동학적 요소를 환경 동역학과 명시적으로 분리하여 장면 과적합을 완화하고 로봇-환경 상호 작용 모델링을 용이하게 합니다. 폐쇄 루프 제어를 달성하기 위해, 고주파 로봇 운동 제어와 통합된 자기 회귀 비디오 예측 모델을 구축하여 로봇 정책과 인간 원격 조작 모두와의 상호 작용을 가능하게 합니다. 실험 결과, 당사 모델은 제한된 고정 장면 데이터만으로 학습되었음에도 불구하고 뛰어난 동일 도메인 성능과 매우 다양하고 새로운 환경에 대한 강력한 제로샷 일반화 능력을 보여줍니다. GeniWorld는 다운스트림 애플리케이션에서 환경 변화에도 안정적인 정책 평가를 위한 확장 가능한 도구로 사용될 수 있습니다. 또한, 제한된 실제 세계 시연만으로도 GeniWorld는 세계 모델 내에서 다양한 조작 경로를 생성하여 복잡한 환경에서의 다운스트림 정책 성능과 견고성을 향상시킵니다.
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, but they often suffer from limited action controllability and poor generalization to out-of-distribution (OOD) scenarios. To this end, we present GeniWorld, an interactive world model for robots that generalizes robustly across unseen scenarios. Building on pretrained video generative models, we use URDF-based rendering to transform numerical actions into visual action representations, enabling spatially grounded action control. By explicitly decoupling embodiment kinematics from environmental dynamics, our model mitigates scene overfitting and facilitates modeling of robot-environment interactions. To achieve closed-loop control, we construct an autoregressive video prediction model integrated with high-frequency robot kinematic control, enabling interaction with both robot policies and human teleoperators. In our experiments, even when trained solely on limited fixed-scene data, our model achieves superior in-domain performance and robust zero-shot generalization to highly randomized, unseen environments. For downstream applications, GeniWorld serves as a scalable policy evaluator that remains reliable under environmental perturbations. Furthermore, even with limited real-world demonstrations, GeniWorld generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.