2608.04830v1 Aug 05, 2026 cs.AI

ContextWeave: 실제 업무 워크플로우 벤치마크

ContextWeave: A Real-World Workflow Benchmark

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xinchi Chen
Xinchi Chen
Citations: 1,454
h-index: 17
Enyu Zhou
Enyu Zhou
Citations: 2,998
h-index: 13
Zhikai Lei
Zhikai Lei
Citations: 464
h-index: 4
Hang Yan
Hang Yan
Citations: 247
h-index: 7
Shichun Liu
Shichun Liu
Citations: 943
h-index: 10
Honglin Guo
Honglin Guo
Fudan University
Citations: 554
h-index: 10
Lifeng Ji
Lifeng Ji
Citations: 18
h-index: 2
Xipeng Qiu
Xipeng Qiu
Citations: 60
h-index: 4
Tianyu Huai
Tianyu Huai
Citations: 168
h-index: 6
Enxi Wang
Enxi Wang
Citations: 0
h-index: 0
Luozhijie Jin
Luozhijie Jin
Citations: 54
h-index: 3
Yang Liu
Yang Liu
Citations: 0
h-index: 0
Y. Suo
Y. Suo
Citations: 14
h-index: 1
Lizhi Lin
Lizhi Lin
Citations: 111
h-index: 5
Pengfang Qian
Pengfang Qian
Citations: 143
h-index: 3

언어 모델 에이전트가 독립적인 작업에서 장기적이고 상태를 유지하는 워크플로우로 발전함에 따라 메모리는 필수적이지만, 기존의 평가 방법은 종종 이를 검색 또는 질의 응답으로 단순화합니다. 본 연구에서는 실제 사무 환경에서의 워크플로우 성능 향상에 기억된 경험이 얼마나 기여하는지를 평가하는 장기적인 벤치마크인 ContextWeave를 소개합니다. ContextWeave는 14명의 사용자의 개인 정보를 보호하여 재구성한, 최대 3개월 분량의 워크플로우를 1,005개의 실행 가능한 작업으로 구성하며, 여기에는 지침, 컨테이너화된 환경, 경로 및 작업별 평가 기준이 포함된 568개의 핵심 평가 작업이 있습니다. 이 벤치마크는 작업 공간의 품질과 사용자 개인별 선호도와의 일관성을 측정하며, 관련성, 연속성, 해결 가능성 및 오해를 불러일으킬 수 있는 정보에 대한 강건성 진단 기능을 제공합니다. 고정된 모델 하에서 6가지 메모리 구성 요소를 비교한 결과, 가장 효과적인 구성은 작업 공간 점수를 68.08에서 78.20으로, 선호도 점수를 41.50에서 70.60으로 향상시켰습니다. 고정된 메모리 구성 요소 하에서도, 테스트된 모든 5개의 기본 모델에서 검색 기능은 두 가지 모두의 성능을 향상시키지만, 그 정도는 모델에 따라 크게 다릅니다. 분석 결과, 실행 가능한 경험 기반 메모리는 간결한 요약보다 워크플로우 연속성을 유지하고 불필요한 탐색을 줄이는 데 더 효과적이지만, 오해를 불러일으킬 수 있는 정보에 더 취약할 수도 있습니다. 이러한 연구 결과를 바탕으로, 검색의 관련성뿐만 아니라 실행 중에도 신뢰성 있게 사용될 수 있도록 메모리 시스템을 최적화하는 것이 중요합니다.

Original Abstract

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!