HumanStudy-Bench: 참여자 시뮬레이션을 위한 AI 에이전트 설계
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
거대언어모델(LLM)은 사회과학 실험에서 가상 참여자로 점점 더 많이 사용되고 있지만, 그들의 행동은 종종 불안정하며 설계 방식에 매우 민감하다. 기존 평가들은 흔히 기반 모델의 능력과 실험적 구현을 혼동하여, 결과가 모델 자체를 반영하는지 아니면 에이전트 설정을 반영하는지 불분명하게 만든다. 대신 우리는 참여자 시뮬레이션을 전체 실험 프로토콜에 걸친 에이전트 설계 문제로 정의하며, 여기서 에이전트는 기반 모델과 행동 가정을 담고 있는 명세(예: 참여자 속성)로 정의된다. 우리는 LLM 기반 에이전트를 조율하여 '필터-추출-실행-평가' 파이프라인을 통해 출판된 인간 대상 실험을 재구성하는 벤치마크이자 실행 엔진인 HUMANSTUDY-BENCH를 소개한다. 이는 시행 시퀀스를 재현하고 원래의 통계 절차를 종단간(end-to-end)으로 보존하는 공유 런타임 내에서 원본 분석 파이프라인을 실행한다. 과학적 추론 수준에서의 충실도를 평가하기 위해, 우리는 인간과 에이전트의 행동이 얼마나 일치하는지 정량화하는 새로운 지표를 제안한다. 우리는 개인 인지, 전략적 상호작용, 사회 심리학을 아우르는 12개의 기초 연구를 이 동적 벤치마크의 초기 세트로 구현하였으며, 이는 수십 명에서 2,100명 이상의 참여자에 이르는 인간 표본을 포함한 6,000회 이상의 시행을 포괄한다.
Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base-model capabilities with experimental instantiation, obscuring whether outcomes reflect the model itself or the agent setup. We instead frame participant simulation as an agent-design problem over full experimental protocols, where an agent is defined by a base model and a specification (e.g., participant attributes) that encodes behavioral assumptions. We introduce HUMANSTUDY-BENCH, a benchmark and execution engine that orchestrates LLM-based agents to reconstruct published human-subject experiments via a Filter--Extract--Execute--Evaluate pipeline, replaying trial sequences and running the original analysis pipeline in a shared runtime that preserves the original statistical procedures end to end. To evaluate fidelity at the level of scientific inference, we propose new metrics to quantify how much human and agent behaviors agree. We instantiate 12 foundational studies as an initial suite in this dynamic benchmark, spanning individual cognition, strategic interaction, and social psychology, and covering more than 6,000 trials with human samples ranging from tens to over 2,100 participants.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.