2605.29893v1 May 28, 2026 cs.AI

중복인가 필수인가? 에이전트 경로에서 중복 단계를 탐지하기 위한 벤치마크

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

Minyang Hu
Minyang Hu
Citations: 28
h-index: 3
Zhinuo Zhou
Zhinuo Zhou
Citations: 10
h-index: 1
Jiacheng Liang
Jiacheng Liang
Citations: 72
h-index: 3
Jiahao Guo
Jiahao Guo
Citations: 7
h-index: 1
Yiyang Yin
Yiyang Yin
Citations: 15
h-index: 2
Bo Yang
Bo Yang
Citations: 1
h-index: 1
Xiongwei Han
Xiongwei Han
Citations: 14
h-index: 2

LLM 기반 에이전트는 다단계 추론과 도구 사용을 통해 복잡한 작업을 해결하는 데 강력한 능력을 보여주었습니다. 그러나 기존의 평가 프로토콜은 주로 작업 성공 여부에 초점을 맞추고, 에이전트 행동의 중요한 측면인 실행 효율성을 간과합니다. 실제로 에이전트 경로에는 작업 완료에 거의 기여하지 않으면서 상당한 자원을 소비하는 중복 단계가 자주 포함됩니다. 본 연구에서는 에이전트 경로에 대한 새로운 연구 분야인 **중복 단계 탐지**를 제안하고 정립합니다. 이를 지원하기 위해, 본 연구에서 제시하는 **RedundancyBench**는 다양한 작업과 신중하게 주석 처리된 경로를 포함하는 새로운 벤치마크이며, 각 단계는 작업 완료에 대한 기여도에 따라 레이블이 지정됩니다. RedundancyBench를 사용하여, 경로 내의 단계가 중복인지 필수적인지를 판단하기 위한 세 가지 대표적인 방법을 개발하고 평가했습니다. 실험 결과, 가장 성능이 좋은 방법조차도 중복 단계를 탐지하는 데 24.88%의 낮은 정확도를 보였으며, 일부 방법은 무작위 추측보다 더 나쁜 성능을 나타냈습니다. 이러한 결과는 이 작업의 복잡성을 강조하며, 이 분야에 대한 추가 연구의 필요성을 시사합니다. {코드 및 데이터셋은 본 논문에 사용되었으며, 다음 링크에서 확인할 수 있습니다: https://anonymous.4open.science/r/RedundancyBench}.

Original Abstract

LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation protocols primarily focus on task success, overlooking a critical aspect of agent behavior: execution efficiency. In practice, agent trajectories often contain redundant steps that consume substantial resources while contributing little to task completion. In this work, we propose and formulate a new research area: \textbf{redundant step detection} for agent trajectories. To support this initiative, we introduce \textbf{RedundancyBench}, a new benchmark that contains diverse tasks with carefully annotated trajectories, where each step is labeled according to its contribution to task completion. Using RedundancyBench, we develop and evaluate 3 representative methods to answer whether a step within trajectory is redundant or necessary. Our results show that even the best-performing method achieves only 24.88\% score in detecting redundant steps, while some methods perform worse than random guessing. These results highlight the task's complexity and the need for further research in this area. \footnote{Code and dataset in this paper are both available in \href{https://anonymous.4open.science/r/RedundancyBench}{https://anonymous.4open.science/r/RedundancyBench}.}

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!