EgoGenesis: 온라인 앵커 프로젝티브 메모리와 액션-3D RoPE를 활용한 개인 중심 세계-행동 모델링
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
개인 시점(egocentric) 영상은 임베디드 AI에게 풍부한 조작 경험을 제공하지만, 다양한 장면, 객체, 동작 및 몸체의 데이터를 수집하는 것은 여전히 비용이 많이 듭니다. 본 논문에서는 EgoGenesis( extit{method})라는 개인 시점 세계-행동 시뮬레이터를 제안합니다. EgoGenesis는 제어 가능하고 고품질의 조작 영상을 생성하여 부족한 실제 데이터 학습을 확장합니다. EgoGenesis는 사전 훈련된 비디오 생성 모델을 기반으로 하며, 두 가지 기하학적 정보를 고려한 조건부 메커니즘을 도입합니다. 온라인 앵커 프로젝티브 메모리(OAPM)는 자동 회귀 생성을 수행하는 동안 첫 번째 프레임의 3D 장면 정보를 유지하고 주기적으로 최근 상태를 업데이트합니다. 액션-3D 로터리 포지션 임베딩(A3D-RoPE)은 카메라 정보를 고려한 3D 로터리 좌표로 엔드 이펙터 동작을 인코딩하여, 골격-영상 교차 어텐션에 동작의 기하학적 정보를 주입하고 정밀한 제어를 가능하게 합니다. 이러한 구성 요소들은 긴 개인 시점 영상 생성에서 시각적 충실도, 기하학적 안정성 및 동작 정렬을 향상시킵니다. 또한, 400개의 실제 트랙터리에 EgoGenesis로 생성된 400개의 트랙터리를 추가함으로써 단일 팔 작업의 실시간 로봇 성공률이 77%에서 84%로, 그리고 양팔 작업의 성공률이 53%에서 70%로 향상되었으며, 이는 합성 데이터가 다운스트림 WAM(World-Action Modeling) 일반화 성능을 크게 향상시킨다는 것을 보여줍니다.
Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.