2602.18940v1 Feb 21, 2026 cs.AI

DREAM: 에이전트 기반 지표를 활용한 심층 연구 평가

DREAM: Deep Research Evaluation with Agentic Metrics

E. Avraham
E. Avraham
Citations: 117
h-index: 4
Changhao Li
Changhao Li
Citations: 28
h-index: 3
R. Dorfman
R. Dorfman
Citations: 2
h-index: 1
Roy Ganz
Roy Ganz
Citations: 432
h-index: 13
Oren Nuriel
Oren Nuriel
Citations: 276
h-index: 8
Amir Dudai
Amir Dudai
Citations: 187
h-index: 7
Aviad Aberdam
Aviad Aberdam
Citations: 608
h-index: 14
Noah R. Flynn
Noah R. Flynn
Citations: 224
h-index: 9
Elman Mansimov
Elman Mansimov
Citations: 5,249
h-index: 15
Aditya Kalyanpur
Aditya Kalyanpur
Citations: 32
h-index: 2
Ron Litman
Ron Litman
Amazon
Citations: 643
h-index: 11

심층 연구 에이전트는 분석가 수준의 보고서를 생성하지만, 단일한 정답(ground truth)의 부재와 연구 품질의 다차원적 특성으로 인해 이를 평가하는 것은 여전히 까다롭다. 최근의 벤치마크들은 각기 다른 방법론을 제안하고 있으나, 표면적으로 뛰어난 유창성과 인용의 일치성이 기저에 있는 사실 및 추론의 결함을 가릴 수 있는 '합성의 신기루(Mirage of Synthesis)' 문제를 겪고 있다. 우리는 치명적인 역량 불일치를 드러내는 4개의 수직적 영역에 걸친 분류 체계를 도입하여 이러한 격차의 특징을 규명한다. 즉, 정적 평가자는 시간적 타당성과 사실적 정확성을 평가하는 데 필요한 도구 사용 능력이 본질적으로 부족하다. 이를 해결하기 위해, 우리는 평가 자체를 에이전트 중심(agentic)으로 전환하여 역량 동등성(capability parity)의 원칙을 구현하는 프레임워크인 DREAM(에이전트 기반 지표를 활용한 심층 연구 평가)을 제안한다. DREAM은 질의 독립적(query-agnostic) 지표와 도구 호출 에이전트가 생성한 적응형(adaptive) 지표를 결합한 평가 프로토콜을 통해 평가를 구조화하며, 이를 통해 시간 인지적 커버리지, 근거 기반 검증 및 체계적인 추론 탐색을 가능하게 한다. 통제된 평가 결과에 따르면, DREAM은 기존 벤치마크에 비해 사실 및 시간적 저하(decay)에 훨씬 더 민감하게 반응하며, 확장 가능하고 참조가 필요 없는(reference-free) 평가 패러다임을 제공한다.

Original Abstract

Deep Research Agents generate analyst-grade reports, yet evaluating them remains challenging due to the absence of a single ground truth and the multidimensional nature of research quality. Recent benchmarks propose distinct methodologies, yet they suffer from the Mirage of Synthesis, where strong surface-level fluency and citation alignment can obscure underlying factual and reasoning defects. We characterize this gap by introducing a taxonomy across four verticals that exposes a critical capability mismatch: static evaluators inherently lack the tool-use capabilities required to assess temporal validity and factual correctness. To address this, we propose DREAM (Deep Research Evaluation with Agentic Metrics), a framework that instantiates the principle of capability parity by making evaluation itself agentic. DREAM structures assessment through an evaluation protocol combining query-agnostic metrics with adaptive metrics generated by a tool-calling agent, enabling temporally aware coverage, grounded verification, and systematic reasoning probes. Controlled evaluations demonstrate DREAM is significantly more sensitive to factual and temporal decay than existing benchmarks, offering a scalable, reference-free evaluation paradigm.

3 Citations
0 Influential
7.5 Altmetric
40.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!