2608.02580v1 Aug 03, 2026 cs.RO

Ego2Robot: 인간의 시점 데이터를 활용한 로봇 데이터 합성 - 확장 가능한 방법

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Zixing Lei
Zixing Lei
Citations: 723
h-index: 6
Tong Zhang
Tong Zhang
Citations: 41
h-index: 4
Zhixuan Liang
Zhixuan Liang
Citations: 797
h-index: 10
Haoqi Yuan
Haoqi Yuan
Citations: 551
h-index: 12
Haoyang Li
Haoyang Li
Citations: 848
h-index: 7
An-Jen Chen
An-Jen Chen
Citations: 52
h-index: 4
Xiong-hui Chen
Xiong-hui Chen
Citations: 799
h-index: 4
Jie Zhang
Jie Zhang
Citations: 12
h-index: 2
Ziyuan Jiao
Ziyuan Jiao
Citations: 24
h-index: 3
Chenxi Xiao
Chenxi Xiao
Citations: 22
h-index: 3
Ye Wang
Ye Wang
Citations: 116
h-index: 6
Peibin Lin
Peibin Lin
Citations: 5
h-index: 1
Yiyang Huang
Yiyang Huang
Citations: 68
h-index: 4
Tao Zhang
Tao Zhang
Citations: 62
h-index: 5
Qin Jin
Qin Jin
Citations: 1,680
h-index: 6

일반화된 로봇 조작 정책을 학습하려면 대규모이고 다양한 데모 데이터가 필요합니다. 인간의 시점에서 촬영된 조작 영상은 풍부한 장면과 작업 다양성을 제공하며, 기존 연구에서는 이러한 영상을 로봇 형식의 데이터로 변환하면 소규모 환경에서 각 작업에 효과적인 정책을 얻을 수 있음을 보여주었습니다. 그러나 이 방법이 비전-언어-행동 모델에 대한 사전 훈련 효과를 대규모로 제공할 수 있는지 여부는 아직 탐구되지 않았습니다. 본 논문에서는 **Ego2Robot**이라는 확장 가능한 파이프라인을 제시합니다. Ego2Robot은 인간의 시점에서 촬영된 조작 영상을 액션 재매핑, 로봇 팔 시각 합성, 다단계 품질 검증 과정을 거쳐 로봇 훈련 데이터로 변환합니다. Ego2Robot은 선별된 데이터셋뿐만 아니라 실제 환경에서 수집된 영상도 지원하며, 15가지 다양한 로봇 형태에 대한 18,561시간 분량의 로봇 훈련 데이터를 생성하여 현재까지 가장 큰 인간-로봇 데이터셋을 구축했습니다. 일반화 성능 평가를 위해 RoboTwin2.0을 확장하여 시각적 외관, 장면 레이아웃, 로봇 형태, 작업 의미 등 다양한 측면에서 독립적인 변경(perturbation)을 적용할 수 있도록 했습니다. 실험 결과, Ego2Robot으로 생성된 데이터와 실제 로봇 데이터를 함께 사용하여 사전 훈련하면 다양한 종류의 변경에 대한 일반화 성능이 꾸준히 향상되는 것을 확인했으며, 이는 실제 로봇 환경에서도 검증되었습니다. 프로젝트 페이지: https://www-ye.github.io/ego2robot_blog/

Original Abstract

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present \textbf{Ego2Robot}, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!