2607.24744v1 Jul 27, 2026 cs.RO

구체화된 조작을 위한 데이터 피라미드

Data Pyramid for Embodied Manipulation: A Survey

Tianxing Chen
Tianxing Chen
Citations: 677
h-index: 10
Yao Mu
Yao Mu
Citations: 294
h-index: 4
Xiaowei Chi
Xiaowei Chi
Citations: 610
h-index: 14
Kuangzhi Ge
Kuangzhi Ge
Citations: 77
h-index: 5
Sirui Han
Sirui Han
Citations: 41
h-index: 4
Shanghang Zhang
Shanghang Zhang
Citations: 387
h-index: 10
Y. Lou
Y. Lou
Citations: 9
h-index: 2
Lingdong Kong
Lingdong Kong
Citations: 249
h-index: 9
Yingshuo Wang
Yingshuo Wang
Citations: 6
h-index: 1
Weihao Yuan
Weihao Yuan
Citations: 1,140
h-index: 15
Ping Luo
Ping Luo
Citations: 1,863
h-index: 14
Ziwei Liu
Ziwei Liu
Citations: 28
h-index: 1
Yifan Ye
Yifan Ye
Citations: 10
h-index: 2
Yankai Fu
Yankai Fu
Citations: 366
h-index: 7
Ya-hui Lv
Ya-hui Lv
Citations: 41
h-index: 3
Bohan Hou
Bohan Hou
Citations: 37
h-index: 3
Jun Cen
Jun Cen
Citations: 73
h-index: 4
Duo Zheng
Duo Zheng
Citations: 631
h-index: 10
Jiaming Liu
Jiaming Liu
Citations: 889
h-index: 14
Ziang Cao
Ziang Cao
Citations: 1,358
h-index: 15
Wei Chow
Wei Chow
Citations: 404
h-index: 8
Xian Sun
Xian Sun
Citations: 74
h-index: 2
Xidong Zhang
Xidong Zhang
Citations: 14
h-index: 2
Zhibo Pang
Zhibo Pang
Citations: 27
h-index: 3
Yiwu Zhong
Yiwu Zhong
Citations: 280
h-index: 3
Zhihe Lu
Zhihe Lu
Citations: 9
h-index: 2
Qifeng Chen
Qifeng Chen
Citations: 136
h-index: 7
Michael Yu Wang
Michael Yu Wang
Citations: 7
h-index: 2
Jianfei Yang
Jianfei Yang
Citations: 16
h-index: 1

다중 모달 기반 모델은 인터넷 전체를 학습하여 보고 듣는 능력을 갖추었습니다. 그러나 구체화된 에이전트는 그러한 단축경로를 사용할 수 없습니다. 왜냐하면 이들은 관찰, 물리적 상태 및 동작을 연결하는 데이터를 필요로 하기 때문입니다. 이러한 정보는 다양한 데이터 소스를 통해 부분적으로 제공될 수 있습니다. 본 연구에서는 실제 로봇 데이터, UMI 스타일 데이터, 개인 중심 및 외부 중심 데이터, 시뮬레이션 데이터, 그리고 일반적인 이미지-텍스트 데이터의 5가지 상호 보완적인 소스로 구성된 '데이터 피라미드'를 통해 구체화된 데이터 생태계를 체계적으로 정리합니다. 우리는 확장성과 로봇 적합성 간의 긴장 관계를 중심으로 이 피라미드를 구성하고, 각 데이터 소스를 데이터 품질, 다양성, 재사용성 및 물리적 충실도 측면에서 특성화합니다. 또한, 최근 개발된 구체화된 기반 모델들의 학습 과정(데이터 레시피)을 분석하여, 사전 훈련 과정에서 다양한 데이터 소스가 어떻게 선택되고 정렬되며 혼합되는지를 조사합니다. 시각-언어-행동 모델, 월드-액션 모델 등 다양한 유형의 구체화된 브레인 모델에 대해, 데이터 구성이 인식, 추론, 계획, 행동 생성 및 세계 예측 능력과 어떤 관련이 있는지 분석합니다. 마지막으로, 대규모 촉각 데이터셋 구축, 실패 및 복구 데이터 수집, 확장 가능한 데이터 수집 파이프라인 개발, 다양한 구체화 방식 간의 동작 정렬, 개인 중심 데이터를 활용한 숙련된 조작, 그리고 로봇 학습을 위한 체계적인 데이터 레시피 설계 등 6가지 주요 과제를 제시합니다. 본 연구가 차세대 구체화 시스템 설계의 기반을 마련하는 데 기여하기를 바랍니다.

Original Abstract

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

1 Citations
0 Influential
7.5 Altmetric
38.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!