MMaDA-VLA: 통합된 다중 모드 명령어 및 생성 기능을 갖춘 대규모 확산 기반 시각-언어-행동 모델
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
시각-언어-행동(VLA) 모델은 시각적 정보와 자연어 명령을 로봇의 행동으로 변환합니다. 그러나 계층적이고 자기 회귀적인 방식은 종종 구조적인 부담을 초래하고, 장기적인 오류를 누적시키며, 환경 동역학을 파악하기 위한 보조 모듈이 필요합니다. 이에 우리는 MMaDA-VLA를 제안하는데, 이는 완전히 독립적이고 사전 훈련된 이산 확산 기반 VLA 모델로, 다중 모드 이해와 생성을 통합합니다. 특히, MMaDA-VLA는 공유된 이산 토큰 공간을 사용하여 미래 목표 관찰과 행동 조각을 동시에 노이즈 제거하여, 보조 세계 모델 없이 예측된 시각적 결과에 기반한 행동을 구현합니다. 이러한 방식으로 병렬적이고 순서에 구애받지 않는 개선은 장기적인 일관성을 향상시킵니다. 광범위한 실험 및 종합적인 분석 결과, MMaDA-VLA는 LIBERO 데이터셋에서 평균 성공률 98.0%, CALVIN 데이터셋에서 평균 성공 시퀀스 길이가 4.78을 달성했으며, 실제 환경에서도 뛰어난 성능을 보였습니다. 프로젝트 페이지는 https://yliu-cs.github.io/MMaDA-VLA 에서 확인할 수 있습니다.
Vision-Language-Action (VLA) models map visual observations and natural-language instructions to robot actions; however, hierarchical and autoregressive paradigms often incur architectural overhead, accumulate long-horizon errors, and require auxiliary modules to capture environment dynamics. To this end, we present MMaDA-VLA, a fully native, pretrained discrete diffusion VLA that unifies multi-modal understanding and generation. Specifically, MMaDA-VLA uses a shared discrete token space to jointly denoise a future goal observation and an action chunk, grounding actions in predicted visual outcomes without an auxiliary world model. In this way, parallel, order-free refinement improves long-horizon consistency. Extensive experiments and comprehensive analyses demonstrate that MMaDA-VLA achieves an average success rate of 98.0\% on LIBERO and an average successful sequence length of 4.78 on CALVIN, while performing strongly in real-world settings. The project page is available at https://yliu-cs.github.io/MMaDA-VLA.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.