2608.03958v1 Aug 04, 2026 cs.AI

기반 모델을 위한 게임 이론: 유사성 추론을 통한 합리적인 협력의 새로운 경로

A game theory for foundation models shows new paths to rational cooperation through similarity inference

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
B. A. Y. Arcas
B. A. Y. Arcas
Citations: 33,584
h-index: 20
Marissa A. Weis
Marissa A. Weis
Citations: 7
h-index: 1
Maciej Wolczyk
Maciej Wolczyk
Citations: 371
h-index: 6
Rajai Nasser
Rajai Nasser
Citations: 8
h-index: 1
R. Saurous
R. Saurous
Citations: 15,619
h-index: 29
João Sacramento
João Sacramento
Citations: 89
h-index: 5
J. Manyika
J. Manyika
Citations: 8,220
h-index: 15
Guillaume Lajoie
Guillaume Lajoie
Citations: 55
h-index: 4
Seijin Kobayashi
Seijin Kobayashi
Citations: 766
h-index: 11
Angelika Steger
Angelika Steger
Citations: 48
h-index: 3
Blake A. Richards
Blake A. Richards
Citations: 192
h-index: 5
Marcus Hutter
Marcus Hutter
Citations: 825
h-index: 5

기반 모델로 구동되는 자율 에이전트가 사회 및 경제 시스템에 점점 더 통합됨에 따라, 이들의 집단 행동을 규제하는 원리를 이해하는 것은 안전과 협력을 보장하는 데 필수적입니다. 의사 합리적인 상호 작용을 모델링하는 주류 프레임워크인 고전 게임 이론은 '분리된 에이전트'라는 가정에 기반합니다. 즉, 에이전트는 자신의 의사 결정을 환경 및 다른 행위자와 독립적으로 간주합니다. 그러나 현대 AI 에이전트는 외부 관찰과 함께 자신의 미래 행동을 동시에 예측합니다. 본 연구에서는 주목할 만한 결과를 보고합니다. 스타일화된 사회적 딜레마 상황에서 최적의 계획을 수행하는 기반 모델 에이전트는 일관되게 안정적인 협력 상태로 수렴하며, 이는 고전 게임 이론의 상호 이기심(mutual defection) 예측과 정면으로 충돌합니다. 이러한 현상을 이해하기 위해, 본 연구에서는 '임베디드 베이지안 에이전트'라는 기반 모델 에이전트를 위한 이론적 모델을 소개합니다. 분리된 에이전트에서 임베디드 에이전트로 전환함으로써, 이 에이전트는 자신이 존재하는 우주의 일부로 자신을 모델링하고, 자신의 의사 결정 알고리즘에 대한 인식론적 불확실성을 유지합니다. 우리는 '임베디드 균형'이라는 새로운 해결 개념을 도입하여 유사성 추론을 통해 임베디드 에이전트가 자신의 계획 과정에서의 고찰을 증거로 활용한다는 것을 보여줍니다. 즉, 협력을 결정하는 것은 유사한 파트너의 유사한 결정을 예측합니다. '임베디드 균형'은 기존의 내쉬 균형을 대체하여 현대 AI 에이전트의 사회적 행동에 대한 기초적인 게임 이론을 제공합니다.

Original Abstract

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!