2606.09748v1 Jun 08, 2026 cs.AI

프로세스 수준 피드백 하에서의 심층 연구 에이전트의 다단계 평가

Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

Hongru Wang
Hongru Wang
The Chinese University of Hong Kong, University of Edinburgh
Citations: 2,565
h-index: 25
Rishabh Sabharwal
Rishabh Sabharwal
Citations: 1
h-index: 1
A. Storkey
A. Storkey
Citations: 16,964
h-index: 48
Jeff Z. Pan
Jeff Z. Pan
Citations: 94
h-index: 5

기존의 심층 연구 에이전트(DRA)에 대한 벤치마크는 단일 결과물만을 평가하며, 중요한 질문인 'DRA가 피드백을 통해 보고서를 개선할 수 있는가?'라는 점을 간과합니다. 이를 조사하기 위해, 우리는 두 가지 피드백 환경 하에서 DRA의 다단계 평가를 수행했습니다. 첫째는 외부 진단 신호 없이 에이전트 스스로 보고서를 수정하는 '자기 성찰'이고, 둘째는 연구 전략상의 부족한 부분을 지적하는 가이드라인을 제공하는 '프로세스 수준 피드백'입니다. 프로세스 수준 피드백을 가능하게 하기 위해, 우리는 '연구 격차 추론(RGI)'이라는 방법을 설계했습니다. 이 방법은 충족 및 미충족된 평가 기준의 패턴을 분석하여 연구 과정상의 격차를 파악합니다. 우리의 분석 결과, 세 가지 주요 사실이 밝혀졌습니다: (i) 자기 성찰 하에서는 에이전트가 평가 기준을 통합하고 이를 다시 퇴화시키는 비율이 거의 동일하여, 전반적인 개선 효과는 미미합니다; (ii) 단 한 번의 프로세스 수준 피드백은 상당한 개선을 가져오며, 정규화된 점수를 약 8~15점 높이고, 대략 35~40%의 기준 통합률을 보입니다; (iii) 이러한 개선 효과는 후속 단계에서 누적되지 않습니다. 왜냐하면 에이전트가 남아 있는 격차를 해결하기 위해 전체 보고서를 다시 작성할 때, 이전에 충족된 최대 24%의 기준에 대해 퇴화 현상을 보이기 때문입니다. 목표 지향적인 가이드라인을 제공하더라도, 평가한 DRA 아키텍처는 여전히 신뢰성 있는 다단계 개선을 달성하기 어렵습니다. 우리의 코드와 결과는 다음 링크에서 공개적으로 이용할 수 있습니다: https://github.com/sabharwalrishabh/Multi-Turn-Evaluation-of-DRAs.

Original Abstract

Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their reports when guided by feedback? To investigate this, we conduct a multi-turn evaluation of DRAs under two feedback settings: self-reflection, in which the agent revises its report without any external diagnostic signal, and process-level feedback, in which the agent receives guidance targeting gaps in its research strategy. To enable process-level feedback, we design Research Gap Inference (RGI), a method that analyzes patterns of satisfied and unsatisfied rubric criteria to infer research-process gaps. Our analysis reveals three key findings: (i) under self-reflection, agents incorporate and regress on rubric criteria at nearly equal rates, yielding negligible net improvement; (ii) a single round of process-level feedback yields substantial gains, raising the normalized score by approximately $8$-$15$ points and yielding a roughly $35$-$40\%$ incorporation rate; (iii) these gains do not compound over subsequent turns, as agents regress on up to $24\%$ of previously satisfied criteria when rewriting the full report to address remaining gaps. Even with targeted guidance, reliable multi-turn improvement remains out of reach for the DRA architectures we evaluate. Our code and results are publicly available at https://github.com/sabharwalrishabh/Multi-Turn-Evaluation-of-DRAs.

0 Citations
0 Influential
43.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!