2604.24558v1 Apr 27, 2026 cs.AI

계층적 행동 공간

Hierarchical Behaviour Spaces

Scott Fujimoto
Scott Fujimoto
Citations: 59
h-index: 4
M. Matthews
M. Matthews
Citations: 182
h-index: 5
A. Kanervisto
A. Kanervisto
Citations: 4,471
h-index: 18
J. Foerster
J. Foerster
Citations: 3,687
h-index: 23
P. D'Oro
P. D'Oro
Citations: 802
h-index: 10
Mikael Henaff
Mikael Henaff
Citations: 173
h-index: 4

최근 계층적 강화 학습 연구에서는 미리 정의된 옵션 보상 함수 집합을 사용하여 학습할 때 수십억 개의 타임스텝에 이르기까지 확장하는 데 성공을 거두었습니다. 본 연구에서는 옵션당 단일 보상 함수를 사용하는 대신, 보상 함수를 효과적으로 활용하여 행동 공간을 유도할 수 있음을 보여줍니다. 제어기는 보상 함수에 대한 선형 조합을 지정할 수 있도록 함으로써, 더욱 풍부한 정책 집합을 표현할 수 있습니다. 이를 계층적 행동 공간(Hierarchical Behaviour Spaces, HBS)이라고 명명합니다. 우리는 NetHack 학습 환경에서 HBS를 평가하여 뛰어난 성능을 입증했습니다. 일련의 실험을 통해, 기존의 통념과는 달리, 본 연구에서 계층 구조의 이점은 장기적인 추론보다는 탐색 능력 향상에서 비롯된다는 것을 확인했습니다.

Original Abstract

Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option reward functions. We show that, instead of using a single reward function per option, the reward functions can be effectively used to induce a space of behaviours, by letting the controller specify linear combinations over reward functions, allowing a more expressive set of policies to be represented. We call this method Hierarchical Behaviour Spaces (HBS). We evaluate HBS on the NetHack Learning Environment, demonstrating strong performance. We conduct a series of experiments and determine that, perhaps going against conventional wisdom, the benefits of hierarchy in our method come from increased exploration rather than long term reasoning.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!