VitaBench 2.0: 장기 사용자 상호 작용에서 개인화 및 능동적인 에이전트 평가
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
대규모 언어 모델(LLM)은 실제 작업에서 사용자와 협력하는 인터랙티브 에이전트로 진화했습니다. 이러한 환경에서의 효과적인 협업은 종종 명시적으로 언급된 정보 외에도 사용자를 이해하는 데 달려 있으며, 사용자 의도는 단편적인 일상 상호 작용에 반영되는 경우가 많으므로 개인 맞춤 모델링과 능동적인 상호 작용이 필요합니다. 그러나 기존 에이전트 벤치마크는 주로 추론 및 도구 사용을 평가하며, 현실적인 시나리오에서 사용자 선호도를 추론하고 활용하는 데 따르는 과제를 간과하는 경향이 있습니다. 이러한 격차를 해소하기 위해, 우리는 장기 사용자 상호 작용에서 개인화되고 능동적인 에이전트의 동작을 평가하기 위한 벤치마크인 VitaBench 2.0을 소개합니다. VitaBench 2.0에서는 작업들이 개별 사용자에 대한 시간 순서대로 구성된 시퀀스로 조직되어 있으며, 사용자 선호도는 단편적이고 이질적인 상호 작용에 내재되어 있습니다. 작업을 성공적으로 완료하려면 에이전트가 이러한 상호 작용으로부터 지속적으로 사용자 선호도를 추출하고 활용하며 업데이트해야 합니다. 또한, 우리는 에이전트가 누락된 정보를 인식하고 의사 결정을 하기 전에 사용자 또는 환경에서 적극적으로 해당 정보를 획득하도록 요구하는 작업들을 통해 능동성을 평가합니다. 체계적인 분석을 지원하기 위해, 다양한 메모리 아키텍처 간의 제어된 비교를 가능하게 하는 확장 가능한 메모리 인터페이스를 제공합니다. 우리는 다양한 최첨단 독점 및 오픈 소스 LLM을 벤치마킹했습니다. 결과는 최첨단 모델에서도 실질적인 개인화가 여전히 매우 어려운 과제라는 것을 보여주며, 현재의 기능과 실제 요구 사항 간에 상당한 격차가 있음을 드러냅니다. 추가적인 분석은 현재 에이전트가 실제 개인 맞춤 의사 결정에서 실패하는 원인과 능력 제한을 밝히고, 향후 모델 개선을 위한 통찰력을 제공합니다.
Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such settings increasingly depends on understanding the user beyond what is explicitly stated, as user intent is often reflected in fragmented daily interactions and requires both personalized modeling and proactive interaction. However, existing agent benchmarks primarily evaluate reasoning and tool use, largely overlooking the challenges of inferring and leveraging user preferences in realistic scenarios. To address this gap, we introduce VitaBench 2.0, a benchmark for evaluating personalized and proactive agent behavior in long-term user interactions. In VitaBench 2.0, tasks are organized as temporally ordered sequences for individual users, where preferences are embedded in fragmented and heterogeneous interactions. Successful completion of tasks requires the agent to continuously extract, utilize, and update user preferences from these interactions. We further evaluate proactiveness through tasks that require agents to recognize missing information and actively acquire it from users or environments before making decisions. To support systematic analysis, we provide an extensible memory interface that enables controlled comparison across different memory architectures. We benchmark a diverse set of frontier proprietary and open-source LLMs. Results show that real-world personalization remains highly challenging even for state-of-the-art models, revealing a substantial gap between current capabilities and practical requirements. Extensive analysis further reveals the failure modes and capability bottlenecks of current agents in real-world personalized decision-making, providing insights for future model improvements.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.