MARL-GPT: 다중 에이전트 강화 학습을 위한 기반 모델
MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning
최근 다중 에이전트 강화 학습(MARL) 분야의 발전은 다양한 도전적인 영역과 환경에서 성공적인 결과를 보여주었지만, 일반적으로 각 작업에 특화된 모델을 필요로 합니다. 본 연구에서는 단일 GPT 기반 모델이 StarCraft Multi-Agent Challenge, Google Research Football 및 POGEMA를 포함한 다양한 MARL 환경과 작업에서 학습하고 뛰어난 성능을 발휘할 수 있도록 하는 일관성 있는 방법을 제안합니다. MARL-GPT는 오프라인 강화 학습을 사용하여 대규모 데이터셋(SMACv2의 경우 4억 개, GRF의 경우 1억 개, POGEMA의 경우 10억 개)으로 학습하고, 작업별 튜닝이 필요 없는 단일 트랜스포머 기반 관찰 인코더를 사용합니다. 실험 결과, MARL-GPT는 테스트된 모든 환경에서 기존의 특화된 모델과 경쟁력 있는 성능을 달성했습니다. 따라서 본 연구 결과는 광범위한(상당히 다른) 다중 에이전트 문제에 대한 다중 작업 트랜스포머 기반 모델을 구축하는 것이 가능하며, 이는 자연어 모델링 분야의 ChatGPT, Llama, Mistral과 같은 기본적인 MARL 모델 개발의 길을 열어줄 수 있음을 시사합니다.
Recent advances in multi-agent reinforcement learning (MARL) have demonstrated success in numerous challenging domains and environments, but typically require specialized models for each task. In this work, we propose a coherent methodology that makes it possible for a single GPT-based model to learn and perform well across diverse MARL environments and tasks, including StarCraft Multi-Agent Challenge, Google Research Football and POGEMA. Our method, MARL-GPT, applies offline reinforcement learning to train at scale on the expert trajectories (400M for SMACv2, 100M for GRF, and 1B for POGEMA) combined with a single transformer-based observation encoder that requires no task-specific tuning. Experiments show that MARL-GPT achieves competitive performance compared to specialized baselines in all tested environments. Thus, our findings suggest that it is, indeed, possible to build a multi-task transformer-based model for a wide variety of (significantly different) multi-agent problems paving the way to the fundamental MARL model (akin to ChatGPT, Llama, Mistral etc. in natural language modeling).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.