SkillMemo: 전문가 지향의 기술 기억 프레임워크를 활용한 복합적인 로봇 조작
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
디퓨전 정책(DP) 및 비전-언어-액션(VLA) 모델을 포함한 로봇 시각-운동 모델은 로봇 조작 벤치마크에서 유망한 성능을 보여주었습니다. 그러나 이러한 모델의 잠재력은 대규모 로봇 동작 데이터셋의 부족으로 인해 근본적으로 제한되며, 이는 분포 외(out-of-distribution) 환경에서의 불충분한 복합 일반화 및 재사용 가능한 기술 구조를 포착하는 능력 부족으로 이어집니다. 이러한 한계를 해결하기 위해, 우리는 장기 시퀀스 데모를 잠재적인 기본 기술로 분해하고, 복합 작업 해결을 위한 동적 에피소드 메모리 뱅크에 기술 수준의 특징을 통합하는 Skill-Based Memory (SkillMemo) 프레임워크를 제안합니다. 구체적으로, 우리는 먼저 Mixture-of-Experts (MoE) 아키텍처를 기반으로 한 전문가 지향 트래젝토리 분할 모듈을 도입하여, 학습된 게이팅 계수를 통해 트래젝토리를 서로 다른 기본 기술로 분할합니다. 또한, 우리는 간결한 기술 표현을 검색 가능한 키-값 쌍으로 저장하는 기술 수준의 에피소드 메모리 아키텍처를 설계했습니다. 추론 과정에서, 메모리 뱅크는 가장 관련성이 높은 기본 기술을 검색하고, 이를 모델의 현재 게이팅 분포와 결합하여 액션 예측을 개선하기 위한 강력한 문맥적 선행 정보를 제공합니다. 시뮬레이션 벤치마크 및 실제 로봇 조작 작업에 대한 광범위한 실험 결과는 SkillMemo가 DP 및 VLA 백본 모두를 지속적으로 향상시키며, 최첨단 성능을 달성하고 $π_{0.5}$보다 우수한 성능을 보이며, 동시에 새로운 작업 구성에 대해 강력한 복합 일반화를 보여준다는 것을 입증합니다.
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model's current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming $π_{0.5}$, while exhibiting strong compositional generalization to unseen task configurations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.