2605.27817v1 May 27, 2026 cs.RO

비디오 모델을 활용한 범용 로봇 제어 정책 개발

Turning Video Models into Generalist Robot Policies

Xingjian Bai
Xingjian Bai
Citations: 490
h-index: 4
Tong Zhao
Tong Zhao
Citations: 87
h-index: 5
Tao Pang
Tao Pang
Citations: 87
h-index: 5
Sizhe Li
Sizhe Li
Citations: 710
h-index: 2
Evan Kim
Evan Kim
Citations: 2
h-index: 1
Max Simchowitz
Max Simchowitz
Citations: 202
h-index: 3
V. Sitzmann
V. Sitzmann
Citations: 466
h-index: 3

비디오 생성 모델은 다양한 형태의 로봇에서 복잡한 작업을 수행하는 모습을 시뮬레이션할 수 있는 유망한 로봇 기반 기술로 부상했습니다. 최근 연구에서는 비디오 모델을 액션 레이블이 포함된 데이터로 미세 조정하여 미래 관측 값과 액션을 동시에 예측하는 로봇 기반 모델을 제안합니다. 본 논문에서는 다른 접근 방식의 한계를 실험합니다. 즉, 비디오 플래너는 그대로 유지하고 로봇의 형태에 특화된 역역학 모델(IDM)을 학습시키는 것입니다. 이러한 분리는 다음과 같은 자연스러운 이점을 제공합니다. 비디오 플래너는 로봇의 형태에 독립적이며, IDM을 재학습하지 않고도 다양한 비디오 모델을 쉽게 교체할 수 있으며, IDM은 사용 가능한 자체 학습 데이터를 통해 독립적으로 학습할 수 있습니다. 본 논문에서는 액션 레이블이 없는 비디오 월드 모델과 로봇의 형태 Jacobian 기반으로 설계된 신중하게 구성된 IDM을 결합한 폐루프 비디오-액션 정책을 제시합니다. 제안하는 IDM 설계는 데이터 효율성이 높고 고차원 액션 공간에도 적용 가능하다는 것을 보여줍니다. 개발된 정책인 Video-to-Embodied Robot Action Model (VERA)은 시뮬레이션 환경과 실제 환경 모두에서 뛰어난 성능을 보이며, 특히 사전 학습 없이 Panda 팔 조작 및 Allegro-hand 로봇을 사용한 정밀한 큐브 회전 작업에서 우수한 성능을 나타냅니다. 동일한 비디오 플래너는 다양한 형태의 로봇에 대해 로봇의 형태에 특화된 IDM과 함께 사용하여 적용할 수 있습니다. 본 연구 결과는 분리된 비디오 계획과 정확한 비디오-액션 변환이 사전 학습 없이도 다양한 형태의 로봇에서 동작하며 일반적인 제어가 가능하도록 하는 실현 가능한 대안임을 보여줍니다. 자세한 내용은 프로젝트 웹사이트(https://vera.csail.mit.edu)에서 확인할 수 있습니다.

Original Abstract

Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly predict future observations and actions by finetuning video models with action-labeled data. In this paper, we test the limits of an alternative approach: leave the video planner as-is while training an embodiment-specific inverse dynamics model (IDM). This decoupling offers several natural benefits: the video planner remains embodiment-agnostic, different video models can be interchanged easily without re-training the IDM, and the IDM can be independently trained with readily available self-play data. We present a closed-loop, video-to-action policy that combines an action-free video world model with a carefully-designed IDM based on the robot embodiment Jacobian. We demonstrate that our IDM design is both data-efficient and scalable to high-dimensional action spaces. Our policy, which we coin the Video-to-Embodied Robot Action Model (VERA), achieves strong performance across simulated and real-world benchmarks, including zero-shot Panda arm manipulation and 16-DoF Allegro-hand dexterous cube re-orientation. The same video planner can be used across multiple embodiments by pairing it with different embodiment-specific IDMs. Our results show that decoupled video planning plus faithful video-to-action translation is a viable alternative route towards zero-shot, cross-embodiment, and generalizable robot control. More results are available on our project website: https://vera.csail.mit.edu.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!