2608.05630v1 Aug 06, 2026 cs.CL

대규모 언어 모델에서의 인간과 유사한 대명사 해소

Human-Like Anaphor Resolution in Large Language Models

Raj Sanjay Shah
Raj Sanjay Shah
Citations: 559
h-index: 12
Keane Zhang
Keane Zhang
Citations: 0
h-index: 0
Varshini Chinta
Varshini Chinta
Citations: 0
h-index: 0
Sashank Varma
Sashank Varma
Citations: 3,081
h-index: 21

대명사는 다른 표현, 즉 선행사를 지칭하는 어절입니다. 이 두 가지를 연결하는 과정을 '대명사 해소'라고 합니다. 인지 과학에서는 담화 구조, 상황 모델 속성 및 의미적 요인 등 대명사 해소의 속도와 성공에 영향을 미치는 다양한 요소들이 밝혀졌습니다. 본 연구에서는 개방형 가중치를 가진 다섯 개의 대규모 언어 모델(LLM)인 GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B 및 Mistral-24B에서 이러한 요소들이 대명사 해소에 미치는 영향을 조사합니다. 처리 난이도를 모델링하기 위해, 우리는 인간의 독해 시간과 대명사에 대한 모델의 놀라움 정도를 연결하는 표준 연결 가설을 채택했습니다. 두 번째 행동적 측정 방법으로, 모델 정확도와 인간의 정확도를 비교하여 대명사의 선행사를 이해하는 능력에 대한 질문에서 얼마나 일치하는지 확인합니다. 결과는 선택적인 인지적 정렬을 보여줍니다. 일부 LLM은 대명사 해소 과정에서 담화의 중요성과 거리 기반 요인에 대해 인간과 유사한 민감성을 보이는 반면, 의미적 간섭 효과에 대해서는 약하거나 전혀 민감성을 보이지 않습니다. 이러한 결과는 LLM이 인간의 대명사 해소를 어느 조건에서 모방하는지를 규정합니다.

Original Abstract

Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!