2608.03044v1 Aug 04, 2026 cs.CL

모방할 것인가, 추정할 것인가? 의견 시뮬레이션을 위한 기초 모델과 추가 학습된 모델의 상반되는 강점

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

Jessica Y. Bo
Jessica Y. Bo
Citations: 115
h-index: 6
Ashton Anderson
Ashton Anderson
Citations: 159
h-index: 5
Difan Jiao
Difan Jiao
Citations: 48
h-index: 3
Seth Grief-Albert
Seth Grief-Albert
Citations: 2
h-index: 1

대규모 언어 모델은 점점 더 인간의 의견을 모방하는 데 사용되고 있지만, 기존 연구에서는 상충되는 결과가 보고됩니다. 일부 연구에서는 인간 설문 데이터와의 유망한 일치성을 발견했지만, 다른 연구에서는 개성 붕괴 및 취약한 인구 통계 민감성이 나타났습니다. 우리는 이러한 갈등의 상당 부분이 두 가지 구별되는 작업에 대한 혼동에서 비롯된다는 것을 보여줍니다. 첫 번째 작업을 '모방(emulation)'이라고 부르는데, 이는 모델이 개별 응답을 생성하여 특정 인구 분포를 형성하는 것입니다. 두 번째 작업을 '추정(estimation)'이라고 부르는데, 이는 모델이 직접적으로 인구 분포를 예측하는 것입니다. Pew American Trends Panel 데이터셋을 사용하여 6개의 일치된 기초 모델과 추가 학습된 모델을 평가한 결과, 기초 모델은 더 강력한 모방 능력을 갖는 것으로 나타났습니다. 즉, 인간의 실제 데이터와 더 유사한 응답 분포를 생성하고 인구 통계 구조를 더 잘 보존했습니다. 반면, 추가 학습된 모델은 직접적으로 분포 예측을 수행할 때 더 정확한 결과를 제공하는 더 강력한 추정 능력을 갖는 것으로 나타났습니다. 우리는 인간 시뮬레이션에 적합한 모델 선택이 텍스트 생성 또는 분포 예측이라는 작업의 요구 사항에 따라 결정되어야 한다고 제안합니다.

Original Abstract

Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We show that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are stronger estimators, producing more accurate distributional predictions when asked directly. We propose that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!