2606.18191v1 Jun 16, 2026 cs.AI

DRFLOW: 개인 맞춤형 워크플로우 예측을 위한 심층 연구 벤치마크

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

I. Laradji
I. Laradji
Citations: 4,198
h-index: 33
Raymond Li
Raymond Li
University of British Columbia
Citations: 243
h-index: 6
Md. Tawkat Islam Khondaker
Md. Tawkat Islam Khondaker
Citations: 571
h-index: 10
Muhammad Abdul-Mageed
Muhammad Abdul-Mageed
Citations: 169
h-index: 8
L. Lakshmanan
L. Lakshmanan
Citations: 19,651
h-index: 73

심층 연구(DR) 시스템은 복잡한 정보 검색 작업에 점점 더 많이 사용되고 있지만, 기존 연구는 주로 보고서 및 요약 생성에 초점을 맞추고 있습니다. 반면, 많은 기업 업무에서는 에이전트가 구체적인 워크플로우를 식별해야 하는데, 이는 일련의 실행 단계를 의미합니다. 예를 들어, 예산 정책을 요약하는 대신, 에이전트는 '정해진 예산 내에서 신규 인력을 요청하려면 어떻게 해야 하는가?'와 같은 질문에 대한 답변을 위해 필요한 단계를 결정할 수 있어야 합니다. 따라서, 본 논문에서는 다양한 소스로부터 예측된 개인 맞춤형 워크플로우를 평가하기 위한 벤치마크인 DRFLOW를 소개합니다. 각 작업에서 에이전트는 흩어져 있는 여러 소스에서 관련 증거를 식별한 다음, 해당 증거를 사용하여 사용자의 작업에 대한 올바른 실행 단계를 예측해야 합니다. DRFLOW는 다섯 가지 영역에 걸쳐 총 100개의 작업을 포함하며, 3,900개 이상의 소스를 기반으로 한 1,246개의 참조 워크플로우 단계를 제공합니다. 우리는 사실적 근거, 단계 복구, 구조적 순서, 조건 해결 및 개인화 측면을 포괄하는 일곱 가지 진단 지표를 정의했습니다. 또한, 개인 맞춤형 워크플로우 예측을 위한 참조 에이전트인 DRFLOW-Agent (DRFA)를 제시합니다. 실험 결과, DRFA는 강력한 기준 에이전트에 비해 성능이 향상되었지만(평균 F1 점수 10.02% 향상), 이러한 워크플로우 지표 측면에서 상당한 개선 여지가 있음을 보여줍니다. 이는 완전하고 정확한 개인 맞춤형 워크플로우를 예측하는 것이 심층 연구 분야에서 여전히 해결해야 할 과제임을 시사합니다.

Original Abstract

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries. In contrast, many enterprise tasks instead require an agent to identify concrete workflows which is a sequence of action-steps. For example, rather than summarizing budgeting policies, an agent should be able to determine the steps needed to answer a question such as: "How do I request new headcount given a fixed budget?". Therefore, we introduce DRFLOW, a benchmark for evaluating personalized workflows predicted by agents from heterogeneous sources. Each task requires the agent to identify relevant evidence from scattered sources, then use that evidence to predict the correct action-step sequence for the user's task. DRFLOW contains 100 tasks across five domains, with 1,246 reference workflow steps grounded in more than 3,900 sources. We define seven diagnostic metrics covering factual grounding, step recovery, structural ordering, condition resolution, and personalization. We further present DRFLOW-Agent (DRFA), a workflow-oriented reference agent to predict personalized workflow. We show that although DRFA improves over strong baseline agents (upto 10.02% average F1 score), there is substantial room for improvement remains across these workflow metrics, indicating that predicting complete and correct personalized workflows remains a challenging frontier for deep research.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!