LAWM-3D: 인간 비디오로부터 3차원 인지 능력을 갖춘 잠재적 행동을 학습하여 일반화 가능한 로봇 세계 모델 구축
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
세계 모델은 에이전트가 실제 환경과의 상호 작용 없이도 예측 및 계획을 수행할 수 있도록 합니다. 그러나, 이러한 세계 모델의 응용은 액션 레이블링 비용이 높고 플랫폼 간에 액션 공간의 이질성이 존재하기 때문에 개방형 환경에서의 로봇 지능 분야에서 제한적입니다. 최근에는 잠재적 액션 모델(LAM)이 비지도 방식으로 인간 비디오로부터 직접 액션 표현을 학습하여 이러한 제약을 완화했습니다. 그러나, 대부분의 기존 LAM은 단일 시점 입력을 사용하며 주로 2차원 픽셀 공간에서 작동합니다. 따라서 다음과 같은 근본적인 질문이 제기됩니다: 다중 시점 비디오를 LAM 학습에 단순히 통합하는 것만으로도 학습된 잠재적 액션에 3차원 인지 능력을 부여할 수 있는가? 본 연구에서는 그 답은 '아니다'임을 보여줍니다. 주된 이유는 미래 프레임의 외관 정보 누수, 카메라 간 외관 차이 및 시점 변화 때문입니다. 이러한 문제를 해결하기 위해 LAWM-3D를 제안합니다. 이는 세 가지 핵심 설계 요소를 결합하여 구성됩니다: (1) 3차원 인지 능력을 갖춘 잠재적 액션을 학습하기 위한 다중 시점 불변 통합 액션 토큰화 방식, (2) 사전 학습된 3차원 기반 모델에 중간 인코더 특징을 연결하는 기하학적 정렬 제약 조건으로, 이를 통해 명시적으로 시야 간의 기하학적 대응 관계를 제공하고, (3) 미래 프레임 외관 정보로부터의 지름길 학습을 방지하고 LAM이 기하학적으로 의미 있는 동작 정보를 중심으로 감독 신호를 집중하도록 하는 비단조 RGB-D 공동 재구성 목표. 중요하게도, 이러한 구성 요소들은 단순히 결합된 것이 아니라 통합적인 목적에 의해 밀접하게 연결되어 있습니다. 대규모 인간 비디오 사전 학습과 로봇 미세 조정을 포함하는 두 단계의 패러다임을 기반으로 한 광범위한 실험 결과는 제안된 3차원 인지 능력을 갖춘 잠재적 액션이 세계 모델 성능을 크게 향상시키며, 생성 품질, 물리적 일관성 및 일반화 능력 측면에서 최첨단(SOTA) 결과를 달성함을 보여줍니다.
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recently, latent action models (LAMs) have alleviated this bottleneck by learning action representations directly from unlabeled human videos in a self-supervised manner. Nevertheless, most existing LAMs rely on single-view inputs and operate primarily in 2D pixel space, raising a fundamental question: can simply incorporating multi-view videos into LAM training endow the learned latent actions with 3D-aware perception? Our study shows that the answer is negative. The primary reasons lie in future-frame appearance leakage as well as inter-camera appearance discrepancies and viewpoint variations. To address these issues, we propose LAWM-3D, which introduces three tightly coupled key designs: (1) a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions; (2) a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, thereby explicitly providing cross-view geometric correspondences; and (3) a non-injective RGB-D joint reconstruction objective that prevents shortcut learning from future-frame appearance information, forcing the LAM to focus supervision on motion cues with geometric significance. Importantly, these components are not simply stacked but are tightly coupled through a unified motivation. Built upon a two-stage paradigm of large-scale human video pretraining followed by robot fine-tuning, extensive experiments demonstrate that the proposed 3D-aware latent actions significantly improve world model performance, achieving SOTA results in generation quality, physical consistency, and generalization ability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.