2602.00685v1 Jan 31, 2026 cs.AI

HumanStudy-Bench: 참여자 시뮬레이션을 위한 AI 에이전트 설계

HumanStudy-Bench: Towards AI Agent Design for Participant Simulation

Yiwen Tu
Yiwen Tu
Citations: 3
h-index: 1
Xuan Liu
Xuan Liu
Citations: 11
h-index: 2
Haojian Jin
Haojian Jin
Citations: 13
h-index: 2
HaoYang Shang
HaoYang Shang
Citations: 54
h-index: 3
Xinyang Liu
Xinyang Liu
Citations: 30
h-index: 3
Yunze Xiao
Yunze Xiao
Citations: 22
h-index: 2
Zizhan Liu
Zizhan Liu
Citations: 39
h-index: 2

거대언어모델(LLM)은 사회과학 실험에서 가상 참여자로 점점 더 많이 사용되고 있지만, 그들의 행동은 종종 불안정하며 설계 방식에 매우 민감하다. 기존 평가들은 흔히 기반 모델의 능력과 실험적 구현을 혼동하여, 결과가 모델 자체를 반영하는지 아니면 에이전트 설정을 반영하는지 불분명하게 만든다. 대신 우리는 참여자 시뮬레이션을 전체 실험 프로토콜에 걸친 에이전트 설계 문제로 정의하며, 여기서 에이전트는 기반 모델과 행동 가정을 담고 있는 명세(예: 참여자 속성)로 정의된다. 우리는 LLM 기반 에이전트를 조율하여 '필터-추출-실행-평가' 파이프라인을 통해 출판된 인간 대상 실험을 재구성하는 벤치마크이자 실행 엔진인 HUMANSTUDY-BENCH를 소개한다. 이는 시행 시퀀스를 재현하고 원래의 통계 절차를 종단간(end-to-end)으로 보존하는 공유 런타임 내에서 원본 분석 파이프라인을 실행한다. 과학적 추론 수준에서의 충실도를 평가하기 위해, 우리는 인간과 에이전트의 행동이 얼마나 일치하는지 정량화하는 새로운 지표를 제안한다. 우리는 개인 인지, 전략적 상호작용, 사회 심리학을 아우르는 12개의 기초 연구를 이 동적 벤치마크의 초기 세트로 구현하였으며, 이는 수십 명에서 2,100명 이상의 참여자에 이르는 인간 표본을 포함한 6,000회 이상의 시행을 포괄한다.

Original Abstract

Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base-model capabilities with experimental instantiation, obscuring whether outcomes reflect the model itself or the agent setup. We instead frame participant simulation as an agent-design problem over full experimental protocols, where an agent is defined by a base model and a specification (e.g., participant attributes) that encodes behavioral assumptions. We introduce HUMANSTUDY-BENCH, a benchmark and execution engine that orchestrates LLM-based agents to reconstruct published human-subject experiments via a Filter--Extract--Execute--Evaluate pipeline, replaying trial sequences and running the original analysis pipeline in a shared runtime that preserves the original statistical procedures end to end. To evaluate fidelity at the level of scientific inference, we propose new metrics to quantify how much human and agent behaviors agree. We instantiate 12 foundational studies as an initial suite in this dynamic benchmark, spanning individual cognition, strategic interaction, and social psychology, and covering more than 6,000 trials with human samples ranging from tens to over 2,100 participants.

3 Citations
0 Influential
1.5 Altmetric
10.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!