2607.11363v1 Jul 13, 2026 cs.CL

샐리-앤 테스트를 넘어: 지식적 쉘링 포인트를 활용한 LLM의 정신 이론 평가

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

G. Keeling
G. Keeling
Citations: 7,329
h-index: 15
Winnie Street
Winnie Street
Citations: 280
h-index: 5
R. Rocca
R. Rocca
Citations: 110
h-index: 5
Sami Boukortt
Sami Boukortt
Citations: 0
h-index: 0

대규모 언어 모델(LLM)에서의 정신 이론(ToM) 평가는 종종 샐리-앤 과제와 유사한 인지 테스트를 포함하며, 이러한 테스트는 사전 학습 과정에서 관련된 유사한 작업에 노출되었기 때문에 쉽게 우회될 수 있으며, 모델의 기능적 ToM 능력을 자연스러운 환경으로 일반화할 수 있는 방식으로 명확하게 평가하지 못합니다. 이러한 문제점을 해결하기 위해, 우리는 견고하고 일반화 가능한 ToM 능력을 측정하도록 설계된 2인 대화 게임인 지식적 비대칭 쉘링 과제(Epistemic Asymmetry Schelling Task, EAST)를 소개합니다. LLM-LLM 쌍이 다양한 수준의 지식 투명성 상태에서 의미론적 쉘링 포인트에 독립적으로 도달하도록 하여, 모델이 ToM을 안정적으로 적용하여 협력을 달성할 수 있는지 평가합니다. 우리의 결과는 기능적인 사회적 추론 능력의 상당한 격차를 보여주며, 최첨단 모델만이 과제의 다양한 지식적 요구 사항을 성공적으로 처리하는 것으로 나타났습니다. 추론 과정 분석 결과, 협력 실패는 주로 사적인 지식을 상호 지식으로 혼동하는 것과 같은 지식 추적 오류로 인해 발생하는 것으로 나타났습니다. 기존의 정적인 벤치마크에서 높은 성능을 보이는 반면, 우리의 연구는 견고한 사회적 추론 능력과 지식 추적이 여전히 중요한 과제임을 보여주며, 이는 향후 LLM 평가 및 개발을 위한 구체적인 목표를 제시합니다.

Original Abstract

Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in ways that generalize to naturalistic settings. To address these issues, we introduce the Epistemic Asymmetry Schelling Task (EAST), a two-player dialogue game designed to benchmark robust and generalizable ToM abilities. By requiring LLM-LLM dyads to independently converge on semantic Schelling points under varying states of epistemic transparency, we evaluate whether models can robustly apply ToM to achieve coordination. Our results reveal a significant capability gap in functional social reasoning, with only frontier models successfully navigating the varying epistemic demands of the tasks. Analysis of reasoning traces shows that coordination failures are primarily driven by epistemic tracking errors, such as conflating private knowledge with mutual knowledge. Despite high performance on traditional static benchmarks, our study shows that robust social reasoning and epistemic tracking remain a critical bottleneck, providing concrete targets for future LLM evaluation and development.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!