LiLa-WAM: 로봇 조작을 위한 경량 잠재 추론 기반 환경-행동 모델
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
환경-행동 모델링은 로봇 제어 분야에서 유망한 패러다임으로 부상했으며, 이는 모델이 관찰에 반응하는 것뿐만 아니라 장면의 변화를 예측할 수 있도록 합니다. 그러나 기존의 환경-행동 모델(WAM)은 종종 상당한 계산 부담을 초래합니다. 픽셀 공간 기반 방법은 제어에 직접적으로 관련 없는 시각적 세부 정보에 많은 자원을 할당하고, 일부 잠재 공간 기반 방법은 추론 공간을 구축하기 위해 다단계 학습이 필요합니다. 이러한 결과로 발생하는 학습 비용은 제한된 계산 예산 하에서 이러한 방법을 훈련하기 어렵게 만듭니다. 본 연구에서는 LiLa-WAM이라는 경량 환경-행동 모델을 제안합니다. LiLa-WAM은 작은 잠재 공간에서 미래를 추론하며, 단일 24GB GPU에서 전체적으로 훈련할 수 있습니다. 이 모델의 핵심 설계는 향후 상태 예측과 행동 생성에 의해 공동으로 형성된 작은 잠재 추론 공간이며, 이는 모델을 경량화하면서도 제어와 잘 정렬되도록 합니다. 또한 작업 명세를 위해, 시각적 특징 공간에서 각 작업을 방향으로 인코딩하는 언어 독립적인 작업 표현인 Visual Transition Token(VTT)을 추가로 제안합니다. RoboTwin~2.0, LIBERO 및 실제 로봇 작업에서의 실험 결과는 LiLa-WAM의 효과를 입증하며, 단일 GPU 훈련으로 50개의 RoboTwin 작업에서 90.48%의 성공률을 달성했습니다.
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.