DailyReport: 일상적인 검색 작업에 대한 검색 에이전트 평가를 위한 개방형 벤치마크
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
검색 에이전트(SA)는 일반적으로 대규모 언어 모델(LLM)을 활용하여 웹 소스를 자율적으로 탐색하고 정보를 종합적인 답변으로 합성하는 복잡한 정보 검색 작업을 지원합니다. 기존의 SA 평가 벤치마크는 주로 실제 사용자 시나리오에서 발생할 가능성이 낮은 특수 작업에 초점을 맞추고 있으며, 거친 수준의 작업 분류 기준에 의존하여 평가 해석력을 제한하는 경향이 있습니다. 이러한 격차를 해소하기 위해, 우리는 일상적인 검색 작업에 대한 SA 역량을 평가하기 위한 개방형 벤치마크인 DailyReport를 소개합니다. 이 벤치마크는 실제 사용자의 광범위하게 논의되고 시의적절한 정보 요구 사항을 반영하는 150개의 개방형 작업과 3,546개의 관련 평가 기준으로 구성되어 있습니다. 각 작업은 하위 작업으로 분해되고 독립적인 차원을 기준으로 계층화된 평가 기준을 사용하여 평가됩니다. 계층적 성능 분석 및 사용자 중심 집계를 통해, 우리는 각 차원에 대한 매우 해석 가능한 점수와 함께 사용자 선호도 점수를 도출했습니다. 17개의 에이전트 시스템에 대한 우리의 결과는 현재 시스템이 여전히 사용자의 기대 수준에 미치지 못한다는 것을 보여줍니다. 향후 연구를 지원하기 위해, 당사의 데이터셋과 코드는 https://github.com/AGI-Eval-Official/DailyReport 에서 공개적으로 이용할 수 있습니다.
Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses. For SAs evaluation, prior benchmarks mainly focus on specialized tasks that are unlikely to arise in real-world user scenarios. Moreover, their reliance on coarse task-level rubrics often limits evaluation interpretability. To bridge this gap, we introduce DailyReport, an open-ended benchmark to evaluate SA capabilities on daily search tasks. It contains 150 open-ended tasks with 3,546 associated rubrics, capturing widely discussed and timely information demands of real-world users. Each task is decomposed into subtasks and evaluated with cascade rubrics across disentangled dimensions. Through cascade performance attribution and user-centric aggregation, we derive highly interpretable scores for each dimension, along with a user preference score. Our results on 17 agentic systems show that current systems still fall short of users' expectations. To facilitate future research, our dataset and code are made publicly available at https://github.com/AGI-Eval-Official/DailyReport.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.