2607.13705v1 Jul 15, 2026 cs.AI

AgentCompass: 에이전트 능력을 평가하기 위한 통합 인프라

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Zichen Ding
Zichen Ding
Citations: 1,785
h-index: 11
Zerun Ma
Zerun Ma
Citations: 529
h-index: 6
Zehao Li
Zehao Li
Citations: 1,039
h-index: 6
Shudong Liu
Shudong Liu
Citations: 544
h-index: 9
Bowen Yang
Bowen Yang
University of Science and Technology of China
Citations: 1,032
h-index: 5
Zhenyu Wu
Zhenyu Wu
Citations: 1,521
h-index: 9
Zun Wang
Zun Wang
Citations: 18
h-index: 2
Qi Zhang
Qi Zhang
Citations: 101
h-index: 4
Zonglin Li
Zonglin Li
Citations: 156
h-index: 2
Shufan Jiang
Shufan Jiang
Citations: 0
h-index: 0
Songyang Zhang
Songyang Zhang
Citations: 116
h-index: 3
Dongsheng Zhu
Dongsheng Zhu
Citations: 94
h-index: 3
Peiheng Zhou
Peiheng Zhou
Citations: 141
h-index: 4
Mocheng Li
Mocheng Li
Citations: 0
h-index: 0
Jiaye Ge
Jiaye Ge
Citations: 3,509
h-index: 7
Kai Chen
Kai Chen
Citations: 205
h-index: 4
Tiaohao Liang
Tiaohao Liang
Citations: 0
h-index: 0
Zixing Shang
Zixing Shang
Citations: 0
h-index: 0
Wenhui Tian
Wenhui Tian
Citations: 0
h-index: 0
Jun Xu
Jun Xu
Citations: 7
h-index: 1
Dingbo Yuan
Dingbo Yuan
Citations: 33
h-index: 2
Qingqiu Li
Qingqiu Li
Citations: 260
h-index: 8
Liwei Wu
Liwei Wu
Citations: 460
h-index: 7

대규모 언어 모델(LLM)이 자율 에이전트로 진화함에 따라, 통일된 평가 인프라의 필요성이 점점 중요해지고 있습니다. 그러나 현재의 평가 파이프라인은 여전히 매우 분산되어 있고 밀접하게 연결되어 있어 재현성을 저해하고 불필요한 엔지니어링을 초래합니다. 이러한 문제를 해결하기 위해, 우리는 LLM 기반 에이전트의 평가를 위한 개방형, 경량화되고 확장 가능한 인프라인 AgentCompass를 소개합니다. AgentCompass는 벤치마크, 실행 환경, 그리고 시스템 구성 요소라는 세 가지 독립적인 구성 요소를 중심으로 평가 프로세스를 구성하여, 복잡한 실행 로직을 재구현하지 않고도 유연한 구성을 가능하게 합니다. 또한, AgentCompass는 오류 허용 기능을 갖춘 비동기 런타임과 포상 회피와 같은 미묘한 실패 모드를 투명하게 진단할 수 있는 종합적인 경로 분석 도구를 제공합니다. 20개 이상의 벤치마크를 다섯 가지 능력 차원에서 기본적으로 지원하는 AgentCompass는 연구 커뮤니티에 확장 가능하고 재현 가능한 인프라를 제공하여 에이전트 연구 발전에 기여합니다.

Original Abstract

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!