계층적 행동 공간
Hierarchical Behaviour Spaces
최근 계층적 강화 학습 연구에서는 미리 정의된 옵션 보상 함수 집합을 사용하여 학습할 때 수십억 개의 타임스텝에 이르기까지 확장하는 데 성공을 거두었습니다. 본 연구에서는 옵션당 단일 보상 함수를 사용하는 대신, 보상 함수를 효과적으로 활용하여 행동 공간을 유도할 수 있음을 보여줍니다. 제어기는 보상 함수에 대한 선형 조합을 지정할 수 있도록 함으로써, 더욱 풍부한 정책 집합을 표현할 수 있습니다. 이를 계층적 행동 공간(Hierarchical Behaviour Spaces, HBS)이라고 명명합니다. 우리는 NetHack 학습 환경에서 HBS를 평가하여 뛰어난 성능을 입증했습니다. 일련의 실험을 통해, 기존의 통념과는 달리, 본 연구에서 계층 구조의 이점은 장기적인 추론보다는 탐색 능력 향상에서 비롯된다는 것을 확인했습니다.
Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option reward functions. We show that, instead of using a single reward function per option, the reward functions can be effectively used to induce a space of behaviours, by letting the controller specify linear combinations over reward functions, allowing a more expressive set of policies to be represented. We call this method Hierarchical Behaviour Spaces (HBS). We evaluate HBS on the NetHack Learning Environment, demonstrating strong performance. We conduct a series of experiments and determine that, perhaps going against conventional wisdom, the benefits of hierarchy in our method come from increased exploration rather than long term reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.