마인드게임즈: 다중 에이전트 LLM의 사회적 및 전략적 추론을 평가하기 위한 실시간 아레나
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
대규모 언어 모델(LLM)은 점점 더 많은 상호 작용형 에이전트로 배포되고 있지만, 이들의 확장된 상호 작용에서의 사회적 및 전략적 추론 능력에 대한 이해는 여전히 부족합니다. 기존의 평가는 정적인 시나리오 또는 단일 게임 벤치마크에 의존하며, 이는 실제 다중 에이전트 환경에서 요구되는 지속적이고 다면적인 추론을 포착할 수 없습니다. 우리는 Mindgames라는 다중 게임 아레나 및 LLM 에이전트를 평가하는 플랫폼을 소개합니다. Mindgames는 '정신 이론(theory of mind)'과 관련된 상호 보완적인 추론 요건을 구현하며, 여기에는 숨겨진 정보 하에서의 믿음 귀인, 반복적인 전략적 상호 작용을 통한 상대 모델링, 지식 불균형 하에서의 협력적 추론, 그리고 사회적 추론에서의 지속적인 속임수가 포함됩니다. TextArena를 기반으로 구축된 Mindgames는 통일된 상호 작용 인터페이스, TrueSkill 기반의 등급 시스템, 그리고 네 가지 게임 환경에 대한 전체 경로 로깅 기능을 제공합니다. 우리는 주요 AI 학회에서 개최하는 2025년 대회 사이클을 통해 Mindgames를 구현했으며, 이 대회에서는 76개 팀에서 제출한 944개의 에이전트를 네 가지 게임(Colonel Blotto, Iterated Prisoner's Dilemma, Codenames, Secret Mafia)으로 평가했습니다. 우리의 분석 결과, 에이전트 수준과 평가 수준 모두에서 제한 사항이 드러났습니다. 규칙 준수 문제는 여전히 주요 장애물이며, 최고 성능을 보이는 시스템은 반복적으로 명시적인 구조적 지지 장치에 의존하고, 리더보드의 유효성은 환경에 따라 크게 다릅니다. 특히, 실패가 빈번한 환경에서는 전략적 능력만큼이나 상대방의 오류에 대한 견고성이 보상으로 주어질 수 있으며, Secret Mafia는 이번 사이클에서 이러한 '오류 생존' 편향을 두드러지게 나타냅니다. 우리는 29,571개의 다중 에이전트 게임 데이터를 공개하며, 각 게임은 턴 단위의 관찰, 액션 및 보상 정보를 포함합니다. 또한 MG-Ref라는 결정론적인 오프라인 토너먼트 프로토콜을 함께 제공합니다. 이 프로토콜은 새로운 에이전트를 최고 순위의 저오류 Stage~II 제출물 풀과 비교하며, 분석에 사용된 동일한 오류 귀인 방식을 적용하여 평가합니다.
Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly understood. Existing evaluations rely on static vignettes or single-game benchmarks that cannot capture the sustained, multi-faceted reasoning that real-world multi-agent settings demand. We introduce Mindgames, a multi-game arena and evaluation platform for LLM agents that operationalizes complementary reasoning demands relevant to ``theory of mind'': belief attribution under hidden information, opponent modeling through repeated strategic interaction, cooperative inference under knowledge asymmetries, and sustained deception in social deduction. Built on TextArena, Mindgames provides a unified interaction interface, TrueSkill-based rating, and full trajectory logging across four game environments. We instantiate Mindgames through a 2025 competition cycle hosted at a major AI conference, which assessed 944 submitted agents from 76 teams across four games: Colonel Blotto, Iterated Prisoner's Dilemma, Codenames, and Secret Mafia. Our analysis surfaces both agent-level and evaluation-level limitations: brittle rule adherence remains a major bottleneck, top-performing systems repeatedly rely on explicit structural scaffolding, and leaderboard validity differs sharply across environments. In particular, failure-heavy environments can reward robustness to opponent errors as much as strategic ability, with Secret Mafia exhibiting a pronounced error-survival confound in this cycle. We release a dataset of 29,571 multi-agent games with turn-level observations, actions, and rewards, together with MG-Ref, a deterministic offline tournament protocol that scores new agents against a frozen reference pool of top-ranked, low-error Stage~II submissions under the same error-attribution lens used in this analysis.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.