EchoVLA: 시너지 효과를 내는 선언적 기억을 갖춘 로봇 비전-언어-행동 모델을 이용한 모바일 조작
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
최근 비전-언어-행동(VLA) 모델의 발전으로 인해, 에이전트는 다중 양식의 지시를 해석하고 복잡한 작업을 수행할 수 있게 되었습니다. 그러나 기존 VLA 모델은 주로 단기적인, 테이블 위에서의 조작에 국한되어 있으며, 에이전트가 변화하는 공간적 맥락 하에서 탐색과 조작을 조정해야 하는 모바일 조작에는 필요한 기억 및 추론 능력이 부족합니다. 본 연구에서는 모바일 조작을 위한 기억 기반 VLA 모델인 EchoVLA를 제시합니다. EchoVLA는 인간 뇌에서 영감을 받은 시너지 효과를 내는 선언적 기억을 통합하며, 여기에는 공간-의미 정보를 담은 장면 기억과 다중 양식 맥락 특징과 함께 작업 수준 경험을 저장하는 사건 기억이 포함됩니다. 두 개의 기억은 현재 관찰 결과, 작업 기록 및 지시에 따라 개별적으로 저장, 업데이트 및 검색되며, 검색된 표현은 조잡한(coarse) 및 세밀한(fine-grained) 어텐션을 통해 기본 팔(base-arm) 확산 정책을 안내하는 데 사용됩니다. 대규모 학습을 지원하기 위해, 우리는 또한 MoMani라는 자동화된 벤치마크를 도입합니다. MoMani는 다중 양식 대규모 언어 모델(MLLM) 기반의 계획 및 피드백 주도 개선을 통해 전문가 수준의 경로를 생성하며, 실제 로봇 데모로 보완됩니다. 포괄적인 시뮬레이션 및 실제 환경에서의 결과는 EchoVLA가 전반적인 성능을 크게 향상시킨다는 것을 보여줍니다. 예를 들어, 시뮬레이션에서 조작/탐색 작업에서 0.52의 최고 성공률과 모바일 조작 작업에서 0.31의 최고 성공률을 달성했으며, 이는 강력한 기준 모델인 $π_{0.5}$보다 각각 +0.20 및 +0.11 더 높습니다.
Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly confined to short-horizon, table-top manipulation, lacking the memory and reasoning capability required for mobile manipulation, where agents must coordinate navigation and manipulation under changing spatial contexts. In this work, we present EchoVLA, a memory-aware VLA model for mobile manipulation. EchoVLA incorporates a synergistic declarative memory inspired by the human brain, consisting of a scene memory that maintains a collection of spatial-semantic maps and an episodic memory that stores task-level experiences with multimodal contextual features. The two memories are individually stored, updated, and retrieved based on current observations, task history, and instructions, and their retrieved representations are fused via coarse- and fine-grained attention to guide base-arm diffusion policies. To support large-scale training, we further introduce MoMani, an automated benchmark that generates expert-level trajectories through multimodal large language model (MLLM)-guided planning and feedback-driven refinement, supplemented with real-robot demonstrations. Comprehensive simulated and real-world results demonstrate that EchoVLA substantially improves overall performance, e.g., it achieves the highest success rates of 0.52 on manipulation/navigation tasks and 0.31 on mobile manipulation tasks in simulation, exceeding the strong baseline $π_{0.5}$ by +0.20 and +0.11, respectively.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.