2608.12547v1 Aug 12, 2026 cs.MA

LLM이 내쉬 균형을 능가하는가? 자기 학습 멀티 에이전트 게임에서의 분산된 조정 능력 평가

Do LLMs Beat Nash? Testing Decentralized Coordination in Self-Play Multi-Agent Games

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Gregory Dudek
Gregory Dudek
Citations: 3
h-index: 1

중앙 제어 장치 없이 배포되는 대규모 언어 모델(LLM) 에이전트는 종종 협력을 위해 통신이 필요하다고 여겨집니다. 본 연구에서는 이러한 통신 없이도 어떤 것이 가능한지 질문합니다. 독립적인 동일한 모델 인스턴스가 서로 통신할 수 없을 때, 각 모델은 상대방에 대해 얼마나 잘 추론하여 조정되지 않은 플레이의 표준 게임 이론 기준을 넘어설 수 있을까요? 우리는 13개의 언어 모델로 구성된 벤치마크를 소개합니다. 각 모델은 자신의 상대방이 동일한 모델을 실행하고 있다는 정보만 제공받고, 기본 게임의 내쉬 균형과 비교하여 평가됩니다. 두 명의 플레이어가 참여하는 다양한 유형의 매트릭스 게임(총 7가지 유형)에서, 플레이어당 2~10개의 행동 옵션을 가지는 경우, 두 개의 특정 모델이 일관되게 내쉬 균형 기준을 능가하며, 일부 유형에서는 최적의 공동 결과를 향해 접근합니다. 반면, 대부분의 공개 가중치 모델은 게임 구조에 따라 크게 달라지는 부분적인 이득만을 얻습니다. 팀 기반 게임에서 4명 이상의 교체 가능한 에이전트가 참여하는 경우, 특히 행동 공간이 커질수록 성능이 현저히 저하됩니다. 이는 이원 관계 게임에서 자기 학습을 통해 얻는 이점이 더 큰 멀티 에이전트 팀으로는 이전되지 않는다는 것을 시사합니다.

Original Abstract

Large language model agents deployed without a central controller are often assumed to require communication to coordinate their actions. We ask what remains possible without it: when independent instances of the same model cannot communicate, can they still reason about their counterparts well enough to exceed the standard game-theoretic baseline for uncoordinated play? We introduce a benchmark of one-shot, no-communication games in which each of thirteen language models is told only that its counterparts are running the same model and is evaluated against the Nash equilibrium of the underlying game. In two-player matrix games spanning seven archetypes and two to ten actions per player, two frontier-hosted models consistently exceed their Nash benchmark, approaching the optimal joint outcome in several archetypes, while most open-weight models achieve only partial gains that vary sharply by game structure. Performance degrades substantially in team-based games with four or more interchangeable agents, particularly as the action space grows, suggesting that whatever capability drives self-play gains in dyadic games does not transfer to larger multi-agent teams.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!