2602.18462v1 Feb 06, 2026 cs.CY

페르소나 기반 LLM을 활용한 가상 설문 응답자의 신뢰성 평가

Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents

Lorenzo Cima
Lorenzo Cima
Citations: 143
h-index: 6
Stefano Cresci
Stefano Cresci
Citations: 77
h-index: 3
T. Fagni
T. Fagni
Citations: 1,148
h-index: 17
M. Avvenuti
M. Avvenuti
Citations: 2,786
h-index: 30
Erika Elizabeth Taday Morocho
Erika Elizabeth Taday Morocho
Citations: 3
h-index: 1

페르소나 기반 LLM을 가상 설문 응답자로 활용하는 것은 계산 사회 과학 및 에이전트 기반 시뮬레이션에서 흔히 사용되는 방법입니다. 그러나 다중 속성 페르소나 프롬프트가 LLM의 신뢰성을 향상시키는지, 아니면 왜곡을 초래하는지에 대한 명확한 이해가 부족합니다. 본 연구는 World Values Survey의 미국 미세 데이터셋을 활용하여 이러한 문제에 대한 평가를 제공합니다. 구체적으로, 두 개의 오픈 가중 채팅 모델과 랜덤 추론 모델을 비교하여 70,000건 이상의 응답자-항목 인스턴스에 대해 성능을 평가했습니다. 분석 결과, 페르소나 프롬프트가 전반적으로 설문 조사와의 일관성을 향상시키지 못하며, 오히려 많은 경우 성능을 저하시키는 것으로 나타났습니다. 페르소나 효과는 매우 이질적이며, 대부분의 항목에서는 변화가 미미한 반면, 일부 질문과 소외된 하위 그룹에서는 상당한 왜곡이 발생하는 것으로 확인되었습니다. 이러한 연구 결과는 현재의 페르소나 기반 시뮬레이션 방식의 중요한 부정적인 영향을 강조합니다. 즉, 인구 통계 조건 설정은 오류를 재분배하여 하위 그룹의 정확성을 저해하고, 결과적으로 잘못된 분석을 초래할 위험이 있습니다.

Original Abstract

Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whether multi-attribute persona prompting improves LLM reliability or instead introduces distortions. Here we contribute to this assessment by leveraging a large dataset of U.S. microdata from the World Values Survey. Concretely, we evaluate two open-weight chat models and a random-guesser baseline across more than 70K respondent-item instances. We find that persona prompting does not yield a clear aggregate improvement in survey alignment and, in many cases, significantly degrades performance. Persona effects are highly heterogeneous as most items exhibit minimal change, while a small subset of questions and underrepresented subgroups experience disproportionate distortions. Our findings highlight a key adverse impact of current persona-based simulation practices: demographic conditioning can redistribute error in ways that undermine subgroup fidelity and risk misleading downstream analyses.

4 Citations
0 Influential
15 Altmetric
79.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!