DyPES-VLA: 공유 동역학 사전 지식 학습 및 로봇의 형태에 특화된 제어를 통한 다양한 로봇 플랫폼에서의 조작
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
비전-언어-행동(VLA) 모델은 로봇 조작 분야에서 강력한 패러다임으로 자리 잡았지만, 서로 다른 형태의 로봇을 위한 단일 통합 정책을 학습하는 것은 여전히 해결해야 할 과제입니다. 기존 방법들은 주로 두 가지 한계를 가지고 있습니다. 첫째, 다양한 시각 및 상호 작용 데이터에 내재된 공유 동역학 사전 지식을 충분히 활용하지 못하여, 로봇의 형태 간의 일반화 성능이 제한됩니다. 둘째, 로봇의 형태에 특정한 동작들을 공통 형식으로 변환하기 위해 광범위한 수동 전처리 과정이 필요합니다. 이러한 한계점을 극복하기 위해, 본 연구에서는 공유 동역학 사전 지식과 로봇의 형태에 특화된 제어를 학습하는 교차-형태 VLA 모델인 DyPES-VLA를 제안합니다. 먼저, 교차-형태 데이터를 사용하여 미래 예측 목표를 통해 비전-언어 모델(VLM)을 학습하여 공유 쿼리 표현이 객체의 움직임, 접촉 및 상호 작용으로 인한 장면 변화를 포착하도록 합니다. 둘째, 형태에 특화된 Mixture-of-Experts (MoE) 액션 헤드는 이러한 공유 동역학 사전 지식을 사용하여 각 로봇의 고유한 동작 공간에서 직접 실행 가능한 제어로 변환하며, 이 과정에서 수동적인 방식으로 다양한 동작들을 공통 형식으로 정렬할 필요가 없습니다. 이 헤드는 주의 레이어를 공유하여 일반적인 시간적 행동 구조를 포착하고, 형태에 특화된 피드포워드 전문가들은 각 로봇의 고유한 운동학적 제약 조건과 제어 의미론을 처리합니다. 본 연구에서 개발한 통합 정책은 시뮬레이션 및 실제 환경 평가 모두에서 최첨단 성능을 달성했으며, LIBERO 데이터셋에서 98.0%, RoboCasa-GR1 데이터셋에서 59.25%, RoboTwin~2.0 데이터셋에서 89.02%의 성공률을 기록했습니다.
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.