2608.05799v1 Aug 06, 2026 cs.RO

XEWorld: 액션 기반 세계 모델이 새로운 로봇 하드웨어에 얼마나 잘 일반화될 수 있는가?

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Nianfeng Liu
Nianfeng Liu
Citations: 12
h-index: 2
Yan Huang
Yan Huang
Citations: 39
h-index: 4
Keji He
Keji He
Citations: 460
h-index: 7
Qisen Ma
Qisen Ma
Citations: 41
h-index: 2
Jiabing Yang
Jiabing Yang
Citations: 33
h-index: 3
Ziheng He
Ziheng He
Citations: 25
h-index: 3

액션 기반 세계 모델은 로봇 제어 분야에서 유망한 학습 시뮬레이터이지만, 기존 연구에서는 주로 훈련에 사용된 로봇만을 대상으로 평가하여 물리적 역학을 제대로 반영하는지, 아니면 단순히 시각적 패턴을 암기하는 것인지 파악하기 어렵습니다. 본 논문에서는 모델이 이전에 접하지 못한 로봇의 움직임을 얼마나 정확하게 모방할 수 있는지 평가하기 위해 XEWorld라는 제어된 교차-하드웨어 테스트베드를 소개합니다. XEWorld는 물리적으로 동일한 환경에서 다양한 로봇을 사용하여 세계 모델을 평가함으로써, 하드웨어에 대한 의존성을 분리합니다. 체계적인 분석 결과, 현재의 모델들은 주로 2차원 시각 패턴 매칭 방식으로 작동하며, 일반화 성능은 물리적 운동학적 유사성보다는 시각적 유사성에 의해 결정됩니다. 이러한 한계로 인해, 추상적인 수치 기반 관절 명령을 일관된 시각적 경로로 변환하는 데 어려움을 겪으며, 정적인 초기 상태에서 발생하는 동적인 시각적 변화를 예측하는 데 실패합니다. 결과적으로, 새로운 로봇 하드웨어에 대한 제로샷 일반화는 물리적 특성을 직접적으로 반영하는 픽셀 기반의 액션과 명시적인 공간-시간 정렬 정보가 필요합니다. 심지어 소량의 데이터 적응을 통해 이러한 제로샷 장벽을 극복하더라도, 강제적인 외관 복구 과정은 기존 로봇 하드웨어에 대한 지식을 손상시키는 결과를 초래합니다. 이러한 실패 사례들은 학습된 물리적 역학이 새로운 시각적 특징에 적용될 수 없다는 중요한 단점을 드러냅니다. 따라서 진정한 교차-하드웨어 일반화를 달성하기 위해서는 시각적 외관과 근본적인 물리적 역학을 분리하는 새로운 아키텍처 혁신이 필요합니다.

Original Abstract

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!