2606.20014v1 Jun 18, 2026 cs.LG

다중 에이전트 게임에서의 계층적 제어: LLM 기반 계획 및 강화 학습 실행

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

Alessandro Sestini
Alessandro Sestini
Citations: 148
h-index: 7
Joakim Bergdahl
Joakim Bergdahl
Citations: 359
h-index: 7
Amir Baghi
Amir Baghi
Citations: 1
h-index: 1
Jean-Philippe Barrette-LaPierre
Jean-Philippe Barrette-LaPierre
Citations: 0
h-index: 0
Florian Fuchs
Florian Fuchs
Citations: 0
h-index: 0
Linus Gisslén
Linus Gisslén
Citations: 4
h-index: 1
Jannik Hosch
Jannik Hosch
Citations: 0
h-index: 0
Konrad Tollmar
Konrad Tollmar
Citations: 44
h-index: 2
Iolanda Leite
Iolanda Leite
Citations: 1
h-index: 1

강화 학습(RL)은 순차적인 의사 결정에서 뛰어난 성능을 보이지만, 희소한 보상, 거대한 상태-행동 공간, 그리고 조정된 전략 학습의 어려움으로 인해 복잡한 다중 에이전트 환경으로 확장하는 데는 여전히 어려움이 있습니다. 본 연구에서는 사전 훈련된 대규모 언어 모델(LLM)을 사용하여 팀 전체의 전문적인 강화 학습 스킬 정책 중에서 선택하는 중앙 집중식 전략 컨트롤러 역할을 수행하고, 강화 학습 정책은 반응적인 저수준 실행을 담당하는 계층적 아키텍처를 제안합니다. 이 하이브리드 시스템을 경쟁적인 2대2 King of the Hill 환경에서 행동 트리(BT)와 '단일' RL(스킬 분해 없이 end-to-end 학습) 기준 성능과 비교했습니다. LLM+RL 시스템은 수동으로 설계된 BT와 통계적으로 동등한 수준의 작업 성능을 달성했으며 (46.4% vs 51.5% 승률, p=0.103), 이는 모두 스킬 분해 없이 학습된 단일 RL보다 훨씬 우수한 성능입니다. 사용자 설문조사($n=15$) 결과, 응답자의 60%가 LLM+RL 에이전트를 가장 인간과 유사하게 인식했으며 (p=0.027), 이는 행동적 적응성과 전술적 다양성 때문이라고 밝혔습니다. 이러한 결과는 사전 훈련된 LLM 추론이 사전 훈련된 강화 학습 스킬을 효과적으로 조정하여 경쟁력 있는 다중 에이전트 협력을 달성하고, 수동 규칙 엔지니어링 없이 우수한 현실감(believability)을 제공할 수 있음을 보여줍니다.

Original Abstract

Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments remains challenging due to sparse rewards, large state-action spaces, and the difficulty of learning coordinated strategies. We propose a hierarchical architecture where a pretrained large language model (LLM) acts as a centralized strategic controller that selects among specialized RL skill policies for a team of agents, while RL policies handle reactive low-level execution. We evaluate this hybrid system in a competitive 2v2 King of the Hill environment against behavior tree (BT) and \emph{``Flat''} RL (end-to-end training without skill decomposition) baselines. The LLM+RL system achieves task performance statistically equivalent to hand-crafted BT (46.4\% vs 51.5\% win rate, $p=0.103$) while both significantly outperform Flat RL trained without skill decomposition. A user study ($n=15$) reveals that 60\% of participants perceive LLM+RL agents as the most human-like ($p=0.027$), citing behavioral adaptability and tactical variability. These results demonstrate that pretrained LLM reasoning can effectively orchestrate pretrained RL skills, achieving competitive multi-agent coordination and superior perceived believability without manual rule engineering.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!