단순 질의 응답에서 심층 연구로: 반복적인 작업 진화를 통해 구축된 검증 가능한 벤치마크
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
심층 연구 벤치마크는 전문가 수준의 과제와, 해당 분야 지식을 기반으로 한 신뢰성 있는 평가를 요구합니다. 기존 벤치마크들은 주로 전문가의 직접 작성 또는 사전 제작된 자료에 의존하며, 완전 자동화 방식은 일관성과 추적 가능한 검증을 보장하기 어렵다는 한계가 있습니다. 이러한 문제점을 해결하기 위해, 우리는 31개의 주제와 10개의 주요 범주를 포괄하는 500개의 심층 연구 과제로 구성된 검증 가능한 벤치마크를 개발했습니다. 이 벤치마크는 심층 연구에 필요한 상호 보완적인 능력을 평가하기 위한 세 가지 유형의 질의 방식을 포함합니다. 벤치마크는 간단한 질문을 심층 연구 과제로 점진적으로 변환하는 반복적인 탐색-정형화-도전 과제 파이프라인을 사용하여 자동으로 구축되었습니다. 각 과제는 원자적 단계와 관련된 체크포인트로 구성된 방향성 비순환 그래프(DAG)로 표현되며, 이를 통해 질의, DAG, 평가 기준을 제어된 방식으로 함께 발전시킬 수 있습니다. 실험 결과, 벤치마크는 모델과 질의 유형 간의 명확한 차이를 보여주었으며, 사실 기반의 세부적인 평가 기준은 인간 중심적이고 안정적인 평가를 가능하게 했습니다. 데이터, 구현 코드 및 결과는 공개적으로 제공됩니다.
Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existing human-authored materials, while fully automatic construction struggles to ensure consistent and traceable verification. To address this gap, we introduce a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research. The benchmark is constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline that progressively transforms simple questions into deep research tasks. Each task is represented as a directed acyclic graph (DAG) of atomic steps and associated checkpoints, enabling the query, DAG, and rubrics to evolve together in a controlled manner. Experiments demonstrate that the benchmark clearly discriminates among models and query types, while its fact-grounded pointwise rubrics enable fine-grained, human-aligned, and stable evaluation. Our data, implementation, and results are publicly available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.