2607.14485v1 Jul 16, 2026 cs.AI

사회 시뮬레이션에서 생성형 에이전트를 위한 단계별 선호도 학습

Step-Level Preference Learning for Generative Agents in Social Simulations

Pingyue Sheng
Pingyue Sheng
Citations: 169
h-index: 1
Baicheng Chen
Baicheng Chen
Citations: 7
h-index: 1
Lanlan Qiu
Lanlan Qiu
Citations: 3
h-index: 1
Jian Zhao
Jian Zhao
Citations: 0
h-index: 0
Kang Wang
Kang Wang
Citations: 0
h-index: 0
Yuyang Tian
Yuyang Tian
Citations: 65
h-index: 5
Shunqi Mao
Shunqi Mao
Citations: 71
h-index: 5
Tianxing He
Tianxing He
Citations: 9
h-index: 1
Wenchang Gao
Wenchang Gao
Citations: 2
h-index: 1
Yunfei Ma
Yunfei Ma
Citations: 2
h-index: 1

대규모 언어 모델(LLM) 기반의 생성형 에이전트는 계획, 기억 검색, 성찰 및 행동 선택과 같은 중간 단계를 포함하는 장기적인 의사 결정 과정을 통해 인간의 행동을 시뮬레이션합니다. 그러나 이러한 중간 단계에 대한 세분화된 인간 주석은 여전히 부족하며, 기존 에이전트는 이러한 중간 결정에 대한 인간의 선호도를 반영하지 못합니다. 이러한 격차를 해소하기 위해, 우리는 단계별 인간 선호도 감독을 통해 에이전트 의사 결정 경로에 대한 데이터를 수집할 수 있는 인터랙티브 시뮬레이션 환경인 exttt{method}를 소개합니다. 이를 통해 57,000개의 세분화된 주석 데이터셋을 구축했습니다. 우리는 이 데이터를 사용하여 지도 학습 및 직접 선호도 최적화를 통해 오픈 웨이트 언어 모델에 대한 단계별 선호도 학습을 수행했으며, 시뮬레이션의 정확성, 조정 능력 및 상호 작용 품질을 지속적으로 향상시키고, 에이전트가 더욱 사회적으로 효과적인 행동을 보이도록 유도했습니다. 우리의 결과는 단계별 인간 감독이 로컬 의사 결정 품질과 장기적인 에이전트 행동 모두를 개선하는 데 효과적인 학습 신호라는 것을 보여줍니다.

Original Abstract

Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!