2608.06375v1 Aug 06, 2026 cs.RO

ω-0: 동시 인간형 로봇의 위치 이동 및 조작을 위한 잠재적 예측 세계 행동 모델

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Gen Li
Gen Li
Citations: 11
h-index: 1
Xinying Guo
Xinying Guo
Citations: 1,583
h-index: 7
Jianfei Yang
Jianfei Yang
Citations: 27
h-index: 2
Shanghang Zhang
Shanghang Zhang
Citations: 921
h-index: 16
Xichen Yuan
Xichen Yuan
Citations: 30
h-index: 2
Peiyuan Zhi
Peiyuan Zhi
Citations: 220
h-index: 4
Zhe Li
Zhe Li
Citations: 118
h-index: 7
Zhenzhen Zhang
Zhenzhen Zhang
Citations: 4
h-index: 2
Yangyang Wei
Yangyang Wei
Citations: 35
h-index: 3
Wenjie Zhang
Wenjie Zhang
Citations: 0
h-index: 0
Feng Gao
Feng Gao
Citations: 0
h-index: 0

인간형 로봇이 수행하는 가사 작업은 종종 동시 위치 이동과 조작을 요구하며, 이 과정에서 로봇은 이동, 자세 조정, 균형 유지, 그리고 물체 조작을 하나의 통합된 동작으로 처리해야 합니다. 그러나 기존의 인간형 로봇 제어 방식은 일반적으로 위치 이동과 조작을 분리하여 처리하는 경향이 있으며, 최근 연구에서 개발된 세계-행동 모델은 여전히 팔 중심적이거나 비디오 중심적인 한계를 가지고 있습니다. 본 논문에서는 실제 환경에서의 동시 인간형 로봇 위치 이동 및 조작을 위한 잠재적 예측 전체 신체 행동 모델인 ω-0을 제안합니다. ω-0은 언어 지시, 현재 시각 정보, 그리고 로봇의 자기 인지 상태를 입력으로 받아, 실제 로봇 실행에 적합한 전체 신체 행동의 잠재 변수를 직접적으로 예측합니다. 기존 방식과는 달리, ω-0은 미래 비디오를 재구성하는 대신, 가벼운 예측 목표로서 압축된 미래 시각 정보 표현을 학습하여, 잠재적인 시각적 예측 능력과 확산 기반의 전체 신체 행동 생성 방식을 결합합니다. 본 모델은 로봇 중심 RGB 이미지, 외부 관점 RGB 이미지, 그리고 외부 관점 깊이 정보를 입력으로 사용하며, 제어 기반 시뮬레이션 리플레이를 활용하여 인간/공개 데이터에서 얻은 시각적-운동 선행 지식을 로봇 실행 가능한 행동의 잠재 변수로 연결합니다. 또한, 동기화된 멀티 뷰 관찰 정보, 전체 신체 SMPL 동작, 로봇 상태, 그리고 행동 잠재 변수를 포함하는 40시간 이상의 실제 가사 환경 인간형 로봇 데이터셋인 ω-HOME을 수집했습니다. 11가지 가사 작업에 대한 실제 실험 결과는 단일 ω-0 모델이 부드러운 이동 중 조작 동작을 생성하며, 기존의 모방 학습, VLA, 인간형 로봇 제어, 그리고 세계-행동 모델 기반 방식보다 우수한 성능을 보임을 입증합니다.

Original Abstract

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $ω$-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, $ω$-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, $ω$-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect $ω$-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single $ω$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!