VLA-Pro: 비전-언어-행동 모델을 위한 교차 작업 절차적 기억 전송
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
비전-언어-행동(VLA) 모델은 범용 로봇 조작에 강력한 잠재력을 보여주지만, 여전히 객체, 장면 및 행동 패턴 전체에 걸쳐 관련 경험을 이전해야 하는 새로운 작업으로의 일반화에는 어려움을 겪습니다. 본 논문에서는 VLA-Pro라는 플러그 앤 플레이 프레임워크를 제안합니다. VLA-Pro는 학습 시 작업과 관련된 절차적 기억을 저장하고 추론 중에 이러한 기억을 전송하여 교차 작업 일반화를 향상시키는 것을 목표로 합니다. 구체적으로, VLA-Pro는 학습 과정에서 작업별 LoRA 어댑터를 매개변수화된 절차적 기억으로 저장합니다. 추론 시에, VLA-Pro는 현재의 다중 모달 컨텍스트를 기반으로 관련 절차적 기억을 검색하고 이러한 기억들을 동적으로 결합하여 현재 액션 조각을 생성합니다. RoboTwin, RLBench 및 실제 로봇 조작 작업에서의 실험 결과, VLA-Pro는 여러 백본에서 일관되게 교차 작업 일반화를 향상시키며 시뮬레이션 환경에서는 최대 207%의 상대적인 성능 향상을 달성하고, 실제 환경에서의 성공률을 5.8%에서 65.0%로 증가시켰습니다. 이러한 결과는 절차적 기억 검색 및 적응이 모듈성과 실행 안정성을 유지하면서 새로운 작업에 조작 경험을 이전하는 효과적인 메커니즘임을 시사합니다.
Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task generalization across multiple backbones, achieving up to a 207% relative improvement in simulation and increasing real-world success rate from 5.8% to 65.0%. These results suggest that procedural memory retrieval and adaptation provide an effective mechanism for transferring manipulation experience to novel tasks while preserving modularity and execution stability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.