2607.27816v1 Jul 30, 2026 cs.CL

차용된 역사 너머: 사용자와 연계된 사용자 시뮬레이션을 활용한 인터랙티브 역할극 평가

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

T. China
T. China
Citations: 1,478
h-index: 18
Beijing
Beijing
Citations: 231
h-index: 8
China
China
Citations: 2
h-index: 1
Jie Gao
Jie Gao
Citations: 219
h-index: 3
Yuhan Zhu
Yuhan Zhu
Citations: 0
h-index: 0
Mingxuan Du
Mingxuan Du
Citations: 193
h-index: 3
Benfeng Xu
Benfeng Xu
Citations: 1,349
h-index: 15
Lingyun Yu
Lingyun Yu
Citations: 0
h-index: 0
Hongtao Xie University of Science
Hongtao Xie University of Science
Citations: 0
h-index: 0
MetaStone Technology
MetaStone Technology
Citations: 0
h-index: 0

역할극 에이전트(RPA)는 대규모 언어 모델의 가장 중요한 소비자 응용 분야 중 하나로 자리 잡았습니다. 사용자는 정서적 위안과 같은 경험을 위해 RPA와 다중 턴 대화를 진행하며, 따라서 RPA의 성능을 측정하고 시스템을 비교하며 추가 개선을 위한 지침을 제공하기 위해서는 신뢰성 있는 평가가 필수적입니다. 그러나 기존 벤치마크는 일반적으로 RPA가 고정된 대화 흐름을 이어가도록 강제한 다음, 사용자로부터 분리된 고정된 기준에 따라 결과를 평가합니다. 우리는 이러한 설계 방식의 두 가지 한계를 지적하고 경험적으로 입증했습니다. 첫째, RPA의 출력은 이전 대화 기록에 의해 영향을 받으므로, 실제 다중 턴 환경에서의 역할극 능력을 과학적으로 평가하기 어렵습니다. 둘째, 사용자 경험은 개인마다 크게 다르며, 기존의 고정된 기준이 항상 사용자의 만족도와 일치하지 않을 수 있습니다. 따라서 우리는 사용자와 연계된 언어 모델 기반 역할극 평가 시스템(PALATE: Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation)을 제안합니다. PALATE는 사용자 시뮬레이터를 기반으로 구축된 확장 가능한 RPA 벤치마크입니다. PALATE는 300개의 캐릭터 프로필 풀과 함께 제공됩니다. 주요 평가는 사용자와 관련된 시뮬레이터 5개를 사용하여 후보 RPA가 미리 정의된 캐릭터 프로필을 기반으로 자유로운 다중 턴 대화를 진행하도록 합니다. 일반적인 품질 기준 외에도, 사용자 만족도를 측정하기 위한 개인 맞춤형 기준을 구축했습니다. 홀드아웃 데이터셋에 대한 평가 결과, 개인 맞춤형 기준이 일반적인 기준보다 인간의 판단과 더 높은 일치도를 보였습니다. 16개의 후보 시스템에 대한 주요 평가에서는 각 시스템이 생성한 다중 턴 대화 경로를 통해, RPA의 전반적인 품질, 장기적인 세션 능력 및 사용자별 경험을 개별적으로 분석합니다. 이를 통해 PALATE는 시스템을 단일하고 사용자와 독립적인 순위로 단순화하는 대신, 특정 사용자-RPA 쌍에 대한 해석 가능한 평가 결과를 제공합니다.

Original Abstract

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!