예측 가능하고, 정렬되어 있으며, 확장 가능한 로봇 학습을 향하여
Towards Predictive, Aligned, and Scalable Robot Learning
기본적으로 학습은 단순히 암기를 넘어, 다양한 가능성을 탐색하며 새로운 문제를 해결할 수 있는 능력을 의미합니다. 본 논문에서는 Lumo-2라는 잠재 공간 세계-행동 모델을 소개합니다. 이 모델은 잠재 공간에서 세계의 역학 관계를 추론하여 행동을 생성합니다. 학습된 잠재 공간 세계 역학은 물리적으로 기반한 시각적 변화를 포착하며, 자연스럽게 미래 가능성을 인코딩하고 다양한 모달리티 간의 정렬을 위한 통합적인 기반을 제공합니다. 이러한 구조는 세계 모델링과 유사한 예측적 추론을 가능하게 하면서도, 제어에 관련된 물리적 역학에 집중하여 경량화된 형태를 유지합니다. 본 연구의 핵심 가설은 행동 생성 품질이 잠재 공간의 기하학적 구조에 의해 결정된다는 것입니다. 기존의 재구성 기반 행동 토큰화 방법이 저수준 신호 충실도에 편향된 표현을 유발하여, 재구성 품질과 실제 제어 성능 간의 불일치를 초래한다는 것을 확인했습니다. 이러한 한계를 극복하기 위해, 본 논문에서는 잠재 공간 세계 역학, 시각 정보 및 언어 정보를 통해 행동 표현을 점진적으로 정렬하는 다단계 모달리티 사전 정렬 전략을 제안합니다. 이 과정은 다양한 모달리티 간의 일관성을 강화하고, 추상화를 촉진하며, 예측적 추론을 위한 구조화된 잠재 공간을 유도합니다. 본 연구에서는 잠재 공간 모델링 및 모달리티 정렬에 대한 체계적인 실험 결과를 제시하고, 이러한 요소들이 확장성 법칙과 일반화 성능에 미치는 영향을 분석합니다. 결과는 Lumo-2가 강력한 시각-언어-행동(VLA) 모델 및 세계-행동 모델(WAM) 기준 모델보다 일관되게 우수한 성능을 보이며, 특히 시간적 추론, 물리적 이해 또는 높은 제어 복잡성을 요구하는 어려운 실제 환경 작업에서 더 큰 개선 효과를 나타낸다는 것을 보여줍니다. 이러한 결과는 구조화된 다중 모달리티 정렬 및 예측적 추론이 로봇 지능 발전에 있어 근본적인 원리임을 시사합니다.
Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to world modelling while remaining lightweight and focused on physical dynamics relevant to control. Central to our approach is the hypothesis that action generation quality is governed by the geometry of the latent space. We observe that standard reconstruction-based action tokenization objectives induce representations biased toward low-level signal fidelity, leading to misalignment between reconstruction quality and downstream control performance. To address this limitation, we propose a multi-stage modality pre-alignment strategy in which action representations are progressively aligned with latent world dynamics, vision, and language. This process enforces cross-modal consistency, promotes abstraction, and induces a structured latent space for predictive reasoning. We provide a systematic empirical study of latent world modelling and modality alignment, analyzing their roles in scaling laws and out-of-distribution generalization. Results show that Lumo-2 consistently outperforms strong vision-language-action (VLA) and world-action model (WAM) baselines, with gains on challenging real-world tasks requiring temporal reasoning, physical understanding, or high control complexity, including long-horizon and dexterous manipulation. These findings suggest that structured multimodal alignment and predictive reasoning are fundamental principles for advancing embodied intelligence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.