향상된 샘플 복잡도를 갖는 하이퍼 그라디언트 기반 양층 강화 학습
Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity
양층 강화 학습(Bilevel Reinforcement Learning, BiRL)은 메타 러닝, 계층적 작업 분해 및 인간 피드백 기반 강화 학습(Reinforcement Learning from Human Feedback, RL-HF)과 같은 다양한 문제 유형을 형식화할 수 있는 중요한 강화 학습 프레임워크입니다. 대부분의 양층 강화 학습 알고리즘은 헤세 행렬을 사용하는 하이퍼 그라디언트 때문에 확장성이 떨어지거나, 페널티 기반 근사 방법을 사용하기 때문에 높은 샘플 복잡도를 겪습니다. 본 논문에서는 엔트로피 정규화된 할인 강화 학습 목적 함수에 대한 볼츠만 정책의 최적성을 활용하는 하이퍼 그라디언트 기반 양층 강화 학습 알고리즘을 제안합니다. 제안하는 알고리즘은 헤세 행렬 연산을 사용하지 않으며, 비교적 완만한 규칙성 조건 하에서 반복 복잡도 $O(ε^{-1})$과 최첨단 샘플 복잡도 $ ilde{O}(ε^{-2})$를 달성합니다. 또한, 수렴 분석 과정에서 기존의 최첨단 샘플 복잡도 연구에 포함된 외부 레벨 목적 함수에 대한 Polyak-Lojasiewicz (PL) 조건이라는 가정을 제거할 수 있었습니다.
Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feedback (RL-HF). Most of the bilevel RL algorithms are either not scalable because of using hypergradient with Hessian, or they suffer from high sample complexity because of using penalty-based approximation methods. In this work, we propose a hypergradient-based bilevel RL algorithm using the optimality of the Boltzmann policy for the entropy regularized discounted RL objective function. Our proposed algorithm is Hessian-free and obtains an iteration complexity of $O(ε^{-1})$ and state-of-the-art sample complexity of $\tilde{O}(ε^{-2})$ under mild regularity conditions. Further, in our convergence analysis, we are able to remove the assumption of the Polyak-Lojasiewicz (PL) condition on the outer-level objective function present in the prior state-of-the-art sample complexity work.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.