MobileWAM: 체인-오브-포사이트를 활용한 월드 액션 모델(WAM)과 모바일 매니퓰레이션 간의 연결
MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight
비디오 생성 기반으로 구축된 월드 액션 모델(WAM)은 로봇 학습 분야에서 주목받고 있지만, 주로 테이블탑 조작에 국한되어 있습니다. 반면, 모바일 매니퓰레이션은 장면 전체의 역학적 요소를 고려하면서 동시 이동 및 전신 조작을 요구하지만, 여전히 수동으로 설계된 좌표를 사용하는 시각 인코더가 주를 이루고 있습니다. 본 연구에서는 MobileWAM이라는 트랜스포머 기반 아키텍처를 통해 이러한 간극을 해소하고자 합니다. MobileWAM은 사전 학습된 비디오 디퓨전 트랜스포머와 가벼운 액션 전문가를 레이어별 상호 주의(joint attention) 메커니즘으로 결합하여, 인터넷 규모의 동작 정보를 전신 제어로 변환합니다. 이동 및 조작의 이질적인 역학적 특성을 해결하기 위해, 액션 전문가의 각 피드포워드 레이어는 공유된 전문가, 로케이션 전문가, 그리고 매니퓰레이션 전문가로 구성된 세 가지 전문가를 혼합하여 사용하며, 동작 의도를 기반으로 소프트하게 라우팅합니다. 또한, 감독 학습을 강화하기 위해 Chain-of-Foresight (CoF)라는 방법을 제안합니다. CoF는 중간 표현이 순차적으로 미래의 잠재 상태들을 예측하도록 하며, 각 단계는 이전 단계에 의해 조건화됩니다. CoF는 우리의 분리된 비디오-액션 디노이징 방식과 자연스럽게 결합됩니다. 배포 시, WAM은 현재 프레임을 인코딩하는 역할만 수행하며, 미래 예측 기능은 기울기(gradient)를 통해만 적용되므로, 추론 과정에서 미래 예측 체인 및 비디오 생성 과정은 제거되어 정책 수준의 비용만 남게 됩니다. MobileWAM은 ManiSkill-HAB 데이터셋에서 기존 모바일 매니퓰레이션 모델보다 뛰어난 성능을 보였으며, 다양한 작업 환경에서 강력한 일반화 능력을 보여주는 ARX Lift2 모바일 매니퓰레이터를 대상으로 미세 조정(fine-tuning)되었습니다. 코드 공개는 곧 진행될 예정입니다.
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation amid scene-scale dynamics, yet is still dominated by dynamics-blind visual encoders with hand-crafted coordination. We bridge this gap with MobileWAM, a mixture-of-transformers architecture that fuses a pretrained video diffusion transformer with a lightweight action expert through layerwise joint attention, translating internet-scale motion priors into whole-body control. To reconcile the heterogeneous dynamics of moving and manipulating, each feed-forward layer of the action expert becomes a three-expert mixture of shared, locomotion, and manipulation experts, softly routed by the motion intent in the action tokens. To densify supervision, we further propose Chain-of-Foresight (CoF): intermediate representations sequentially predict a chain of future latent chunks, each step conditioned on its predecessor. CoF pairs naturally with our decoupled video--action denoising scheme. At deployment, the WAM serves as a pure current-frame encoder; foresight acts only through gradients, so at inference the foresight chain and video generation are discarded, leaving only policy-level cost. MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization. Code will be released soon.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.