2604.15695v1 Apr 17, 2026 cs.GT

편집증의 대가: 비정상적인 다중 에이전트 강화 학습 환경에서의 견고한 위험 민감 협력

The Price of Paranoia: Robust Risk-Sensitive Cooperation in Non-Stationary Multi-Agent Reinforcement Learning

D. Ganguly
D. Ganguly
Citations: 3
h-index: 1
C. S. Jonnalagadda
C. S. Jonnalagadda
Citations: 0
h-index: 0
Pratham Chintamani
Pratham Chintamani
Citations: 0
h-index: 0
A. Ananth
A. Ananth
Citations: 1
h-index: 1

협력적 균형은 불안정합니다. 에이전트들이 고정된 환경이 아닌 서로 함께 학습할 때, 학습 과정 자체가 에이전트들이 유지하려는 협력을 불안정하게 만듭니다. 각 에이전트가 수행하는 기울기 업데이트는 파트너가 수행할 행동의 분포를 변화시키고, 협력적 파트너를 확률적 잡음의 원천으로 만듭니다. 특히, 협력 결정에 가장 민감한 부분에서 이러한 현상이 발생합니다. 본 연구에서는 이러한 공동 학습으로 인한 잡음이 조정 게임의 구조를 통해 어떻게 전파되는지 분석하고, 강한 파레토 우위에도 불구하고 표준적인 위험 중립 학습 환경에서 협력적 균형이 지수적으로 불안정하며, 파트너의 잡음이 게임의 임계 협력 수준을 넘어서면 되돌릴 수 없이 붕괴된다는 것을 확인했습니다. 파트너의 불확실성에 대비하기 위해 분포적 견고성을 적용하는 것이 일반적인 대응이지만, 이는 오히려 상황을 악화시킵니다. 위험 회피적인 보상 목표는 이탈 행동에 비해 고분산적인 협력 행동을 더 크게 벌하고, 이는 불안정 영역을 축소하는 대신 확대합니다. 이는 견고성이 적용되는 영역과 불안정성이 발생하는 영역 간의 근본적인 불일치를 보여주는 역설입니다. 우리는 이러한 문제를 해결하기 위해 견고성이 파트너의 불확실성으로 인해 발생하는 정책 기울기 업데이트의 분산을 목표로 해야 하며, 단순히 보상 분포를 목표로 해서는 안 된다는 것을 보여줍니다. 이러한 구분을 통해, 파트너의 예측 불가능성에 대한 온라인 측정값을 기반으로 기울기 업데이트를 조절하는 알고리즘을 개발했습니다. 이 알고리즘은 대칭 조정 게임에서 협력 영역을 확장한다는 것이 증명되었습니다. 본 연구에서는 이러한 접근 방식의 안정성, 샘플 복잡성 및 복지 결과를 통합하기 위해 '편집증의 대가(Price of Paranoia)'를 도입했습니다. 이는 '혼돈의 대가(Price of Anarchy)'의 구조적 이중 개념입니다. '협력 윈도우(Cooperation Window)'와 함께, 이는 학습 알고리즘이 파트너의 잡음 하에서 얼마나 많은 복지를 회복할 수 있는지 정확하게 특성화하며, 균형 안정성과 샘플 효율성 간의 최적 균형을 결정합니다.

Original Abstract

Cooperative equilibria are fragile. When agents learn alongside each other rather than in a fixed environment, the process of learning destabilizes the cooperation they are trying to sustain: every gradient step an agent takes shifts the distribution of actions its partner will play, turning a cooperative partner into a source of stochastic noise precisely where the cooperation decision is most sensitive. We study how this co-learning noise propagates through the structure of coordination games, and find that the cooperative equilibrium, even when strongly Pareto-dominant, is exponentially unstable under standard risk-neutral learning, collapsing irreversibly once partner noise crosses the game's critical cooperation threshold. The natural response to apply distributional robustness to hedge against partner uncertainty makes things strictly worse: risk-averse return objectives penalize the high-variance cooperative action relative to defection, widening the instability region rather than shrinking it, a paradox that reveals a fundamental mismatch between the domains where robustness is applied and instability originates. We resolve this by showing that robustness should target the policy gradient update variance induced by partner uncertainty, not the return distribution. This distinction yields an algorithm whose gradient updates are modulated by an online measure of partner unpredictability, provably expanding the cooperation basin in symmetric coordination games. To unify stability, sample complexity, and welfare consequences of this approach, we introduce the Price of Paranoia as the structural dual of the Price of Anarchy. Together with a novel Cooperation Window, it precisely characterizes how much welfare learning algorithms can recover under partner noise, pinning down the optimal degree of robustness as a closed-form balance between equilibrium stability and sample efficiency.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!