전략, 보상 그 이상: 정형 게임의 행동 기반 임베딩
Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games
전략적 과제를 학습하는 것은 단순히 가르쳐진 내용 이상의 변화를 가져옵니다. 특정 게임에 대한 미세 조정은 다른 게임에서 에이전트의 추론 능력을 향상시키거나 저하시킬 수 있습니다. 그러나 이러한 전략적 능력의 이전 현상을 이해하고 예측하는 것은 대규모 언어 모델(LLM)에게 여전히 중요한 과제입니다. 정형 게임은 명확하게 정의된 보상과 잘 특성화된 균형 행동을 특징으로 하므로, 이 현상을 분석하기에 이상적인 실험 환경을 제공합니다. 본 연구에서는 게임 임베딩이 다양한 게임에 대한 미세 조정 후 LLM의 전략적 능력 변화를 설명하고 예측할 수 있는지 조사합니다. 우리는 냅시 균형의 엔트로피와 상대방의 행동에 대한 최적 반응의 민감성을 포착하는 가벼운 두 가지 특징을 가진 임베딩을 제안합니다. 기존 연구에서 발표된 구조 기반 임베딩은 주로 게임 식별 정보를 암기하며 일반화에는 실패하지만, 우리의 행동 기반 임베딩은 보류 게임에서의 성능 변화를 안정적으로 예측한다는 것을 보여줍니다. 이러한 결과는 LLM의 전략적 능력 이전이 게임의 보상 구조에 의해 결정되는 것이 아니라, 요구되는 의사 결정 행동의 기본적인 구조에 의해 결정된다는 것을 시사합니다.
Learning a strategic task changes more than what is directly taught: fine-tuning on one game can either enhance or degrade an agent's ability to reason in another. Understanding and predicting this transfer of strategic capabilities, however, remains a key challenge for large language models (LLMs). Normal-form games provide an ideal testbed for analyzing this phenomenon, as they feature explicitly defined payoffs and well-characterized equilibrium behaviours. In this work, we investigate whether game embeddings can explain and predict changes in LLM strategic capabilities following fine-tuning across different games. We propose a lightweight two-feature embedding that captures fundamental behavioural demands: the entropy of the Nash equilibrium and the sensitivity of optimal responses to an opponent's action. We show that while existing published structural embeddings primarily memorize game identities and fail to generalize, our behavioural embedding reliably predicts performance changes on held-out games. These results demonstrate that the transfer of strategic capabilities in LLMs is not dictated by the payoff geometry of a game, but by the underlying structure of the decision-making behaviour it requires.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.