2605.29685v1 May 28, 2026 cs.AI

NICE: 사회적 지능을 진단하기 위한 이론 기반 벤치마크

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

Zaifeng Gao
Zaifeng Gao
Citations: 58
h-index: 5
Yixuan Wang
Yixuan Wang
Citations: 12
h-index: 2
Yunjin Qi
Yunjin Qi
Citations: 0
h-index: 0
Zhaojun Jiang
Zhaojun Jiang
Citations: 105
h-index: 5
Hanxi Pan
Hanxi Pan
Citations: 24
h-index: 3
Xiangting Ji
Xiangting Ji
Citations: 11
h-index: 2
Churu Yu
Churu Yu
Citations: 0
h-index: 0
Chunyuan Zheng
Chunyuan Zheng
Citations: 15
h-index: 2
Yingze Chen
Yingze Chen
Citations: 5
h-index: 1
Jie He
Jie He
Citations: 87
h-index: 4
Liuqing Chen
Liuqing Chen
Citations: 17
h-index: 2
Xuan Wu
Xuan Wu
Citations: 8
h-index: 2
Yanfang Liu
Yanfang Liu
Citations: 0
h-index: 0

대규모 언어 모델(LLM)이 감정적 교감 및 고객 서비스와 같은 사회적 맥락에서 점점 더 많이 활용됨에 따라, 인간-AI 상호작용의 품질과 안전성을 위해 LLM의 사회적 지능을 측정하는 것이 중요해졌습니다. 그러나 기존의 사회적 지능 벤치마크는 사회적 능력을 통합된 구조로 구성하는 통일된 프레임워크가 부족하며, 따라서 세분화된 진단을 제공할 수 없습니다. 본 연구에서는 사회 이론에 기반한 최초의 종합적인 진단 평가를 구축하기 위해, 문헌 검토 및 심리 측정 원칙에 따른 다단계 전문가 검증을 통해 사회적 지능 프레임워크를 구성했습니다. 결과적으로 도출된 프레임워크는 4가지 범주와 11가지 차원으로 구성되며, 각 차원은 세분화된 역량 요소로 더욱 구체화됩니다. 이 프레임워크를 바탕으로, 본 연구에서는 대표적인 중국 맥락을 통해 구현된 137개의 항목으로 구성된 진단 벤치마크인 NICE(Norm, Interaction, Cognition, Experience)를 소개합니다. 5가지 최첨단 LLM과 인간 참조 그룹을 대상으로 평가한 결과, 모델들은 전체 정확도 측면에서는 더 높은 점수를 보였지만, 프레임워크가 지적하는 3가지 구체적인 역량 요소(다중 턴 커뮤니케이션, 비언어적 소통, 공감)에서 일관된 약점을 나타냈습니다. 따라서 NICE는 사회적 지능 평가를 이론에 기반한 진단을 통해 LLM의 사회적으로 중요한 약점을 파악하는 방향으로 재정립합니다.

Original Abstract

As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social intelligence has become critical to the quality and safety of human-AI interaction. However, existing social intelligence benchmarks lack a unified framework that organizes social abilities into a unified structure, and therefore cannot enable fine-grained diagnosis. To build the first holistic diagnostic evaluation grounded in social theory, we first construct a social intelligence framework through a literature review and multi-stage expert validation guided by psychometric principles. The resulting framework includes 4 categories and 11 dimensions, each further specified by fine-grained capability facets. Building on this framework, we introduce NICE (Norm, Interaction, Cognition, Experience), a diagnostic benchmark of 137 items operationalized through representative Chinese contexts. Across 5 frontier LLMs and a human reference group, models score higher in aggregate accuracy yet show a consistent weakness in Communication, which the framework localizes to 3 specific capability facets: multi-turn communication, nonverbal communication, and synchrony. NICE thus reframes social intelligence evaluation toward theory-grounded diagnosis of socially consequential weaknesses in LLMs.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!