메타 강화 학습에서의 지식 재활용
Knowledge Reutilization in Meta-Reinforcement Learning
메타 강화 학습은 관련된 작업에서 공유된 구조를 추출하여 빠른 적응을 가능하게 하지만, 기존의 엔드 투 엔드 방법은 종종 작업 추론과 로봇(embodiment)에 특화된 제어를 결합합니다. 이러한 결합은 비모수적인 작업 의미를 가리고, 샘플 효율성을 저하시키며, 에이전트 간 재사용을 제한할 수 있습니다. 본 연구에서는 동역학을 단순화한 에이전트에서 작업 수준의 지식을 학습하고 이를 다양한 로봇에 전달하는 메타-지식 재활용 프레임워크를 제안합니다. 이 프레임워크는 잠재적인 작업 모드를 구성하기 위해 베이지안 비모수적 사전 분포를 사용하며, 작업 수준의 크기 가이드 정보를 생성하기 위한 고차원 정책을 활용합니다. 재사용 가능한 작업 지식을 다양한 로봇에 연결하기 위해, 우리는 의미-크기 인터페이스와 경량화된 시간 적응기를 도입하여, 정적인 메타 지식을 로봇에 특화된 저수준 제어기에 사용할 수 있도록 시간적으로 일관된 부분 목표로 변환합니다. 여러 이동 로봇 에이전트에 대한 실험 결과, 우리 프레임워크는 최첨단 기준 모델과 비교하여 최종 단계 추적 오류를 94.75% ~ 99.79% 줄이고, 약 23.8%의 상호 작용 데이터로 유사한 성능을 달성했습니다.
Meta-reinforcement learning enables fast adaptation by extracting shared structure from related tasks, but existing end-to-end methods often couple task inference with embodiment-specific control. This coupling can obscure non-parametric task semantics, reduce sample efficiency, and limit cross-agent reuse. We propose a meta-knowledge reutilization framework that learns task-level knowledge on a dynamics-simplified agent and transfers it to heterogeneous agents. The framework uses a Bayesian non-parametric prior to organize latent task modes and a high-level policy to generate task-level magnitude guidance. To bridge reusable task knowledge with different embodiments, we introduce a semantic-magnitude interface and a lightweight temporal adaptor, which convert frozen meta-knowledge into temporally aligned subgoals for embodiment-specific low-level controllers. Experiments on multiple locomotion agents show that our framework reduces final-step tracking error by 94.75% -- 99.79% compared with recent state-of-the-art baselines and achieves comparable deployment performance with about 23.8% of their interaction data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.