DynSess: 역할 연기 에이전트를 위한 동적 세션 레벨 평가 및 최적화 프레임워크
DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents
대규모 언어 모델을 활용한 역할 연기는 근본적으로 세션 레벨의 작업이며, 에이전트는 여러 턴에 걸친 대화를 통해 캐릭터 정체성과 상호 작용 품질을 유지해야 합니다. 그러나 기존의 평가 및 최적화 방법은 주로 단일 턴 수준으로 이루어져 장기적인 품질을 제대로 반영하지 못합니다. 본 연구에서는 역할 연기 에이전트를 위한 통합 세션 레벨 프레임워크인 DynSess를 제안합니다. DynSess-Eval은 장기적인 행동에 초점을 맞춘 평가 기준을 사용하여 전체 대화 세션을 평가합니다. 또한, DynSess는 세션 레벨의 보상을 활용하여 다중 턴 예측 검색을 통해 고품질 학습 데이터를 구축하고, DSPO (오프라인) 및 GSRPO (온라인)라는 두 가지 상호 보완적인 방식으로 DynSess-Character를 학습시켰습니다. 실험 결과, DynSess-Eval은 기존 평가 도구보다 인간의 판단과 훨씬 더 일치하는 경향을 보이며, 블라인드 인간 평가에서는 DynSess-Character가 현존하는 가장 강력한 캐릭터 모델에 필적하는 성능을 보이면서도 훨씬 적은 파라미터를 사용하고 강한 역할 일관성과 상호 작용 능력을 유지합니다. 본 연구에서 구축한 데이터셋과 코드는 향후 연구를 촉진하기 위해 공개될 예정입니다.
Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality across extended multi-turn conversations. Yet existing evaluation and optimization methods remain largely turn-level, failing to capture long-horizon quality. We propose DynSess, a unified session-level framework for role-playing agents. DynSess-Eval scores complete dialogue sessions via rubrics targeting long-horizon behaviors. Leveraging its session-level rewards, we construct high-quality training trajectories through multi-turn lookahead search and train DynSess-Character with two complementary variants: DSPO (off-policy) and GSRPO (on-policy). Experiments show that DynSess-Eval aligns with human judgments substantially better than prior evaluators, and blind human evaluation further shows that DynSess-Character matches the strongest character model despite using substantially fewer parameters, while maintaining strong role consistency and interactive ability. Our dataset and code will be released to facilitate future research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.