구체화된 조작을 위한 데이터 피라미드
Data Pyramid for Embodied Manipulation: A Survey
다중 모달 기반 모델은 인터넷 전체를 학습하여 보고 듣는 능력을 갖추었습니다. 그러나 구체화된 에이전트는 그러한 단축경로를 사용할 수 없습니다. 왜냐하면 이들은 관찰, 물리적 상태 및 동작을 연결하는 데이터를 필요로 하기 때문입니다. 이러한 정보는 다양한 데이터 소스를 통해 부분적으로 제공될 수 있습니다. 본 연구에서는 실제 로봇 데이터, UMI 스타일 데이터, 개인 중심 및 외부 중심 데이터, 시뮬레이션 데이터, 그리고 일반적인 이미지-텍스트 데이터의 5가지 상호 보완적인 소스로 구성된 '데이터 피라미드'를 통해 구체화된 데이터 생태계를 체계적으로 정리합니다. 우리는 확장성과 로봇 적합성 간의 긴장 관계를 중심으로 이 피라미드를 구성하고, 각 데이터 소스를 데이터 품질, 다양성, 재사용성 및 물리적 충실도 측면에서 특성화합니다. 또한, 최근 개발된 구체화된 기반 모델들의 학습 과정(데이터 레시피)을 분석하여, 사전 훈련 과정에서 다양한 데이터 소스가 어떻게 선택되고 정렬되며 혼합되는지를 조사합니다. 시각-언어-행동 모델, 월드-액션 모델 등 다양한 유형의 구체화된 브레인 모델에 대해, 데이터 구성이 인식, 추론, 계획, 행동 생성 및 세계 예측 능력과 어떤 관련이 있는지 분석합니다. 마지막으로, 대규모 촉각 데이터셋 구축, 실패 및 복구 데이터 수집, 확장 가능한 데이터 수집 파이프라인 개발, 다양한 구체화 방식 간의 동작 정렬, 개인 중심 데이터를 활용한 숙련된 조작, 그리고 로봇 학습을 위한 체계적인 데이터 레시피 설계 등 6가지 주요 과제를 제시합니다. 본 연구가 차세대 구체화 시스템 설계의 기반을 마련하는 데 기여하기를 바랍니다.
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.