2605.29512v1 May 28, 2026 cs.AI

마인드게임즈: 다중 에이전트 LLM의 사회적 및 전략적 추론을 평가하기 위한 실시간 아레나

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs

Leshem Choshen
Leshem Choshen
Citations: 250
h-index: 9
Yihang Jiang
Yihang Jiang
Citations: 1
h-index: 1
Yoram Bachrach
Yoram Bachrach
Citations: 6,108
h-index: 42
Yuhong Dai
Yuhong Dai
Citations: 38
h-index: 3
Yan-Ru Ju
Yan-Ru Ju
Citations: 15
h-index: 2
Mathieu Laurière
Mathieu Laurière
Citations: 64
h-index: 5
T. Kachman
T. Kachman
Citations: 632
h-index: 9
Ilya Makarov
Ilya Makarov
Citations: 0
h-index: 0
Jianzhu Yao
Jianzhu Yao
Citations: 87
h-index: 4
P. Viswanath
P. Viswanath
Citations: 28,394
h-index: 55
Yitian Huang
Yitian Huang
Citations: 54
h-index: 3
Bobby Cheng
Bobby Cheng
Citations: 21
h-index: 3
Cheston Tan
Cheston Tan
Citations: 75
h-index: 3
I-Chen Wu
I-Chen Wu
Citations: 6
h-index: 1
M. S. Arya
M. S. Arya
Citations: 0
h-index: 0
A. Anish
A. Anish
Citations: 0
h-index: 0
Aditya Ranjan
Aditya Ranjan
Citations: 3
h-index: 1
Yuan Lu
Yuan Lu
Citations: 44
h-index: 3
A. Thoni
A. Thoni
Citations: 0
h-index: 0
Benjamin Kempinski
Benjamin Kempinski
Citations: 15
h-index: 2
Ben Finch
Ben Finch
Citations: 21
h-index: 1
Leon Guertler
Leon Guertler
Citations: 77
h-index: 2
Viraj Nadkarni
Viraj Nadkarni
Citations: 77
h-index: 6
Aliaksei Korshuk
Aliaksei Korshuk
Citations: 82
h-index: 2
Alexander Buyantuev
Alexander Buyantuev
Citations: 1,050
h-index: 16
Siyuan Wu
Siyuan Wu
Citations: 941
h-index: 12
Yu Cheng
Yu Cheng
Citations: 93
h-index: 4
I-Hsuan Chu
I-Hsuan Chu
Citations: 7
h-index: 1
Yu-Yu Yang
Yu-Yu Yang
Citations: 10
h-index: 2
Qi Cao
Qi Cao
Citations: 0
h-index: 0
Yiheng Sun
Yiheng Sun
Citations: 201
h-index: 7
Hongkun Yao
Hongkun Yao
Citations: 154
h-index: 8
Jingxuan Fu
Jingxuan Fu
Citations: 8
h-index: 2
Hao Liao
Hao Liao
Citations: 15
h-index: 2
Mossimo Ebeling
Mossimo Ebeling
Citations: 0
h-index: 0
Govind Arun
Govind Arun
Citations: 30
h-index: 3
Sadhvik Bathini
Sadhvik Bathini
Citations: 4
h-index: 1
K. Phatnani
K. Phatnani
Citations: 11
h-index: 1
Ks Paval
Ks Paval
Citations: 7
h-index: 1
V. Mehta
V. Mehta
Citations: 21
h-index: 1
S. Aravind
S. Aravind
Citations: 21
h-index: 2
Nikhil Arora
Nikhil Arora
Citations: 6
h-index: 1
Tanya Upadhyay
Tanya Upadhyay
Citations: 8
h-index: 1
Amol Bandagale
Amol Bandagale
Citations: 0
h-index: 0
Chun-Pao Hsiao
Chun-Pao Hsiao
Citations: 2
h-index: 1
Yuting Lin
Yuting Lin
Citations: 52
h-index: 4
A. Chung
A. Chung
Citations: 0
h-index: 0
Jeremiah Thomas
Jeremiah Thomas
Citations: 0
h-index: 0
Maria Polukarov
Maria Polukarov
Citations: 4
h-index: 1
Atlas Wang
Atlas Wang
Citations: 52
h-index: 3
K. Wang
K. Wang
Citations: 79
h-index: 5
Tiru Wu
Tiru Wu
Citations: 0
h-index: 0
Jiwei Zhang
Jiwei Zhang
Citations: 4
h-index: 1

대규모 언어 모델(LLM)은 점점 더 많은 상호 작용형 에이전트로 배포되고 있지만, 이들의 확장된 상호 작용에서의 사회적 및 전략적 추론 능력에 대한 이해는 여전히 부족합니다. 기존의 평가는 정적인 시나리오 또는 단일 게임 벤치마크에 의존하며, 이는 실제 다중 에이전트 환경에서 요구되는 지속적이고 다면적인 추론을 포착할 수 없습니다. 우리는 Mindgames라는 다중 게임 아레나 및 LLM 에이전트를 평가하는 플랫폼을 소개합니다. Mindgames는 '정신 이론(theory of mind)'과 관련된 상호 보완적인 추론 요건을 구현하며, 여기에는 숨겨진 정보 하에서의 믿음 귀인, 반복적인 전략적 상호 작용을 통한 상대 모델링, 지식 불균형 하에서의 협력적 추론, 그리고 사회적 추론에서의 지속적인 속임수가 포함됩니다. TextArena를 기반으로 구축된 Mindgames는 통일된 상호 작용 인터페이스, TrueSkill 기반의 등급 시스템, 그리고 네 가지 게임 환경에 대한 전체 경로 로깅 기능을 제공합니다. 우리는 주요 AI 학회에서 개최하는 2025년 대회 사이클을 통해 Mindgames를 구현했으며, 이 대회에서는 76개 팀에서 제출한 944개의 에이전트를 네 가지 게임(Colonel Blotto, Iterated Prisoner's Dilemma, Codenames, Secret Mafia)으로 평가했습니다. 우리의 분석 결과, 에이전트 수준과 평가 수준 모두에서 제한 사항이 드러났습니다. 규칙 준수 문제는 여전히 주요 장애물이며, 최고 성능을 보이는 시스템은 반복적으로 명시적인 구조적 지지 장치에 의존하고, 리더보드의 유효성은 환경에 따라 크게 다릅니다. 특히, 실패가 빈번한 환경에서는 전략적 능력만큼이나 상대방의 오류에 대한 견고성이 보상으로 주어질 수 있으며, Secret Mafia는 이번 사이클에서 이러한 '오류 생존' 편향을 두드러지게 나타냅니다. 우리는 29,571개의 다중 에이전트 게임 데이터를 공개하며, 각 게임은 턴 단위의 관찰, 액션 및 보상 정보를 포함합니다. 또한 MG-Ref라는 결정론적인 오프라인 토너먼트 프로토콜을 함께 제공합니다. 이 프로토콜은 새로운 에이전트를 최고 순위의 저오류 Stage~II 제출물 풀과 비교하며, 분석에 사용된 동일한 오류 귀인 방식을 적용하여 평가합니다.

Original Abstract

Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly understood. Existing evaluations rely on static vignettes or single-game benchmarks that cannot capture the sustained, multi-faceted reasoning that real-world multi-agent settings demand. We introduce Mindgames, a multi-game arena and evaluation platform for LLM agents that operationalizes complementary reasoning demands relevant to ``theory of mind'': belief attribution under hidden information, opponent modeling through repeated strategic interaction, cooperative inference under knowledge asymmetries, and sustained deception in social deduction. Built on TextArena, Mindgames provides a unified interaction interface, TrueSkill-based rating, and full trajectory logging across four game environments. We instantiate Mindgames through a 2025 competition cycle hosted at a major AI conference, which assessed 944 submitted agents from 76 teams across four games: Colonel Blotto, Iterated Prisoner's Dilemma, Codenames, and Secret Mafia. Our analysis surfaces both agent-level and evaluation-level limitations: brittle rule adherence remains a major bottleneck, top-performing systems repeatedly rely on explicit structural scaffolding, and leaderboard validity differs sharply across environments. In particular, failure-heavy environments can reward robustness to opponent errors as much as strategic ability, with Secret Mafia exhibiting a pronounced error-survival confound in this cycle. We release a dataset of 29,571 multi-agent games with turn-level observations, actions, and rewards, together with MG-Ref, a deterministic offline tournament protocol that scores new agents against a frozen reference pool of top-ranked, low-error Stage~II submissions under the same error-attribution lens used in this analysis.

2 Citations
0 Influential
27.5 Altmetric
139.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!