두 명 플레이르 제로섬 게임을 위한 글로벌 정책 공간 응답 오라클
Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
정책 공간 응답 오라클(PSRO) 프레임워크는 심층 강화 학습(DRL)을 사용하여 제한된 전략 집합을 반복적으로 확장함으로써 대규모 제로섬 게임에서의 균형 계산을 확장합니다. 핵심 과제는 제한된 컴퓨팅 예산 하에서 전체 게임을 잘 근사하는 작은 전략 집합을 구축하는 것입니다. 기존 PSRO 방법은 일반적으로 제한된 게임의 보상을 기반으로 계산된 메타 전략에 대한 최적 반응을 사용하여 전략 집합을 확장하는데, 이는 비효율적인 확장을 초래하여 제한적인 글로벌 개선 효과를 가져올 수 있습니다. 우리는 확장 후 전략 집합의 품질을 직접 평가하여 전략 집합 확장을 안내하는 방법을 제안합니다. 특히, Population Exploitability (PE)를 사용하여 제한된 전략 집합이 전체 게임을 얼마나 잘 나타내는지 측정하고, PE를 명시적으로 최소화하는 두 단계의 탐색-선택 프레임워크를 도입합니다. 우리는 이 프레임워크를 Global PSRO라는 실용적인 DRL 기반 알고리즘으로 구현했으며, 이 알고리즘은 후보 응답을 효율적으로 생성하고 파라미터 공유 조건부 신경망을 통해 PE를 추정합니다. 여러 두 명 플레이르 제로섬 게임에서의 실험 결과, Global PSRO는 기존의 PSRO 방법보다 훨씬 적은 정책 반복 횟수로 낮은 exploitability를 달성하며 Nash 균형을 더 잘 근사하는 것으로 나타났습니다.
The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population whose induced game well approximates the full game. Existing PSRO variants typically expand the population using best responses to meta-strategies computed from restricted-game payoffs, which can lead to inefficient expansions that provide limited global improvement. We propose to guide population expansion by directly evaluating the post-expansion population quality. Specifically, we adopt Population Exploitability (PE) to measure how well a restricted strategy set represents the full game, and introduce a two-phase exploration--selection framework that explicitly minimizes PE during expansion. We instantiate this framework as Global PSRO, a practical DRL-based algorithm that efficiently generates candidate responses and estimates PE via parameter-sharing conditional neural networks. Experiments across multiple two-player zero-sum games show that Global PSRO achieves lower exploitability and approximates Nash equilibria with significantly fewer policy iterations than prior PSRO methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.