사회 시뮬레이션에서 생성형 에이전트를 위한 단계별 선호도 학습
Step-Level Preference Learning for Generative Agents in Social Simulations
대규모 언어 모델(LLM) 기반의 생성형 에이전트는 계획, 기억 검색, 성찰 및 행동 선택과 같은 중간 단계를 포함하는 장기적인 의사 결정 과정을 통해 인간의 행동을 시뮬레이션합니다. 그러나 이러한 중간 단계에 대한 세분화된 인간 주석은 여전히 부족하며, 기존 에이전트는 이러한 중간 결정에 대한 인간의 선호도를 반영하지 못합니다. 이러한 격차를 해소하기 위해, 우리는 단계별 인간 선호도 감독을 통해 에이전트 의사 결정 경로에 대한 데이터를 수집할 수 있는 인터랙티브 시뮬레이션 환경인 exttt{method}를 소개합니다. 이를 통해 57,000개의 세분화된 주석 데이터셋을 구축했습니다. 우리는 이 데이터를 사용하여 지도 학습 및 직접 선호도 최적화를 통해 오픈 웨이트 언어 모델에 대한 단계별 선호도 학습을 수행했으며, 시뮬레이션의 정확성, 조정 능력 및 상호 작용 품질을 지속적으로 향상시키고, 에이전트가 더욱 사회적으로 효과적인 행동을 보이도록 유도했습니다. 우리의 결과는 단계별 인간 감독이 로컬 의사 결정 품질과 장기적인 에이전트 행동 모두를 개선하는 데 효과적인 학습 신호라는 것을 보여줍니다.
Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.