2601.19935v1 Jan 13, 2026 cs.CL

Mem2ActBench: 작업 지향형 자율 에이전트의 장기 메모리 활용성 평가를 위한 벤치마크

Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents

Yiting Shen
Yiting Shen
Citations: 39
h-index: 3
Kun Li
Kun Li
Citations: 153
h-index: 8
Wei Zhou
Wei Zhou
Citations: 105
h-index: 6
Songlin Hu
Songlin Hu
Citations: 197
h-index: 8

대규모 언어 모델(LLM) 기반 에이전트는 점점 더 복잡하고 도구 기반의 작업에 활용되고 있으며, 이러한 작업에서는 장기 메모리가 행동을 결정하는 데 중요한 역할을 합니다. 그러나 기존 벤치마크는 주로 에이전트가 명시적인 질문에 대해 개별 사실을 수동적으로 검색하는 능력을 테스트하는 데 초점을 맞추고 있습니다. 이러한 벤치마크는 에이전트가 메모리를 적극적으로 활용하여 작업을 수행하는 능력, 즉 더욱 중요한 능력을 평가하지 못합니다. 이러한 격차를 해소하기 위해, 우리는 에이전트가 적절한 도구를 선택하고 매개변수를 연결하여 장기 메모리를 활용하여 도구 기반 작업을 수행할 수 있는지 평가하는 벤치마크인 extsc{Mem2ActBench}를 소개합니다. 이 벤치마크는 지속적인 어시스턴트 사용 시나리오를 시뮬레이션하며, 사용자가 긴 시간 동안 중단된 상호 작용에서 동일한 주제를 언급하고, 이전에 설정된 선호도와 작업 상태가 암시적으로 적용될 것으로 기대합니다. 우리는 다양한 소스(ToolACE, BFCL, Oasst1)를 결합하고, 일관성 모델을 통해 충돌을 해결하며, 2,029개의 세션을 생성하는 자동화된 파이프라인을 구축했습니다. 각 세션은 평균 12개의 사용자-어시스턴트-도구 턴으로 구성됩니다. 이러한 메모리 체인을 기반으로, 역 생성 방법을 사용하여 400개의 도구 사용 작업을 생성했으며, 인간 평가 결과 91.3%가 장기 메모리에 크게 의존하는 것으로 확인되었습니다. 7개의 메모리 프레임워크에 대한 실험 결과, 현재 시스템은 매개변수 연결을 위해 메모리를 적극적으로 활용하는 데 여전히 부족하며, 이는 작업 실행에서 메모리 활용성을 평가하고 개선하기 위한 더욱 효과적인 접근 방식의 필요성을 강조합니다.

Original Abstract

Large Language Model (LLM)-based agents are increasingly deployed for complex, tool-based tasks where long-term memory is critical to driving actions. Existing benchmarks, however, primarily test a angent's ability to passively retrieve isolated facts in response to explicit questions. They fail to evaluate the more crucial capability of actively applying memory to execute tasks. To address this gap, we introduce \textsc{Mem2ActBench}, a benchmark for evaluating whether agents can proactively leverage long-term memory to execute tool-based actions by selecting appropriate tools and grounding their parameters. The benchmark simulates persistent assistant usage, where users mention the same topic across long, interrupted interactions and expect previously established preferences and task states to be implicitly applied. We build the dataset with an automated pipeline that merges heterogeneous sources (ToolACE, BFCL, Oasst1), resolves conflicts via consistency modeling, and synthesizes 2,029 sessions with 12 user--assistant--tool turns on average. From these memory chains, a reverse-generation method produces 400 tool-use tasks, with human evaluation confirming 91.3\% are strongly memory-dependent. Experiments on seven memory frameworks show that current systems remain inadequate at actively utilizing memory for parameter grounding, highlighting the need for more effective approaches to evaluate and improve memory application in task execution.

13 Citations
1 Influential
4 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!