2608.02046v1 Aug 03, 2026 cs.CL

CompanionBench: 이론 기반의 실제 세계 데이터 기반 AI 정서적 동반자 평가 지표

CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

Junchen Wan
Junchen Wan
Citations: 91
h-index: 4
Jihao Huang
Jihao Huang
Citations: 40
h-index: 3
Yaopei Liu
Yaopei Liu
Citations: 121
h-index: 2
Yumin Huang
Yumin Huang
Citations: 21
h-index: 2
Lei Wang
Lei Wang
Citations: 78
h-index: 3
Guangjia Chai
Guangjia Chai
Citations: 0
h-index: 0

대규모 언어 모델(LLM) 기반의 동반자는 개인에게 중요한 영역에서 널리 사용되고 있지만, 그 성능은 제대로 평가되지 못하고 있습니다. 기존 평가 지표는 사람이 직접 작성한 시나리오와 프롬프트 기반 시뮬레이터를 사용하며, 공감 능력을 하나의 점수로 통합하고, 평가자의 편향(예: 가족 간의 선호, 척도 변화)을 고려하지 않습니다. 본 연구에서는 상호작용형 이중 언어 평가 지표인 CompanionBench를 제안합니다. 저희가 알기로는, CompanionBench는 시나리오와 사용자 시뮬레이터를 모두 익명화된 실제 데이터를 기반으로 구축한 최초의 동반자 평가 지표입니다. 숨겨진 정보 공개 기능은 각 페르소나의 대화 경로를 에이전트의 행동에 따라 분기시켜, 미리 작성된 대본 없이도 상호작용 상태 공간을 제어합니다. 심리학 및 상담 분야의 25개 이론에서 파생된 10가지 역량을 정의했으며, 이 중 4가지는 이전 연구에서 명시적으로 평가되지 않았습니다(예: 모호성 처리, 자기 대상 반응성, 긍정적 공명, 적절한 도전). 에이전트는 주관적인 10가지 역량 평가 기준과 더 깊은 정보 공개를 달성했는지에 대한 결정적 측정이라는 두 가지 상호 보완적인 축을 기준으로 평가됩니다. 서로 다른 배경의 평가자 패널을 구성하여 가족 간의 편향을 줄이고, Item Response Theory 모델을 사용하여 에이전트 품질과 평가자의 엄격함을 분리합니다. 이론은 무엇을 측정할지, 그리고 페르소나가 어떻게 구조화될지를 결정하며, 실제 데이터는 이벤트, 역사 및 프로필을 제공하여 이론적 기반과 현실성을 모두 확보합니다. 순위는 두 언어 모두에서 재현 가능합니다(ρ = 0.996 ZH / 0.953 EN). 28개의 에이전트를 평가한 결과, 전체 점수로 가려졌던 역량 수준의 차이가 드러났습니다. 감정 조절과 적절한 도전은 여전히 흔히 부족한 부분이며, 모호성 처리는 가장 큰 차이를 보여줍니다. 역할극 에이전트는 순위에서 하위에 속했으며, 몰입도가 관계 능력으로 이어지지 않는다는 것을 시사합니다. 전반적으로 에이전트의 주요 실패 원인은 실질적인 관계 지원 대신 표면적인 친밀감을 제공하는 것입니다. 저희는 500쌍의 중국어-영어 병렬 데이터와 평가 코드를 공개할 예정입니다.

Original Abstract

LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!