ACE-Data-0: 인간 중심의 환경 데이터 획득 시스템 - 몸체화된 데이터 엔진
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
몸체화된 지능은 근본적인 데이터 병목 현상에 직면해 있습니다. 모델은 인간이 목표를 추구하는 과정에서 첫인칭 시각, 전신 움직임, 정교한 조작, 객체의 상태, 소리 및 촉감이 어떻게 함께 변화하는지를 포착해야 합니다. 기존의 데이터 세트는 이러한 경험을 다양한 관점, 모달리티 또는 공간 규모로 분할하여 전체적인 인지-행동 루프가 부분적으로만 관찰되도록 합니다. 본 논문에서는 인간 중심의 데이터 엔진인 Ambient Capture Engine (ACE)을 소개합니다. ACE는 실제 가정 환경을 공간적으로 보정되고 시간적으로 동기화된 녹음 스튜디오로 변환합니다. ACE는 상호 보완적인 두 가지 규모에서 작동합니다. 테이블 크기의 구성은 손-객체 조작을 해결하고, 방 크기의 구성은 전신 움직임, 이동 및 가구 배치된 가정 환경에서의 상호 작용을 캡처합니다. ACE는 개인 중심 및 다중 시점 외부 시각 동영상, 전신 및 관절 손 운동, 객체의 기하학적 정보 및 6자유도(DoF) 경로, 오디오 및 촉각 신호를 통합된 다감각 스트림으로 기록합니다. ACE를 사용하여 ACE-Data-0을 구축했으며, 이는 50명의 참가자가 2개의 환경에서 수행한 200가지 작업 범주에 대한 150시간 분량의 17백만 프레임 동영상 데이터를 포함하며, 총 75,000건의 상호 작용 에피소드로 구성됩니다. 이 데이터 세트는 기본적인 조작, 장기적인 가정 활동 연쇄 및 인간-환경 상호 작용을 포괄하며, 단계별 지침 대신 목표 수준에서 자연스러운 행동 변화를 유지합니다. 또한, 신호에서 장면 구성 요소, 그리고 상호 작용으로 점진적으로 발전하는 계층적 벤치마크를 소개합니다. 최첨단 방법론에 대한 평가 결과, 접촉, 가려짐, 자가 움직임 및 장기적인 시간 범위에서 상당한 성능 격차가 존재하는 것을 확인했습니다. ACE-Data-0은 동기화된 인간 시연 데이터를 제공하며, 이는 인지적, 운동학적 및 접촉 정보가 정렬되어 있어 모방 학습, 월드 모델, 비전-언어-액션 시스템 및 몸체화된 AI를 위한 확장 가능한 기반을 제공합니다.
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.