AgentCompass: 에이전트 능력을 평가하기 위한 통합 인프라
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
대규모 언어 모델(LLM)이 자율 에이전트로 진화함에 따라, 통일된 평가 인프라의 필요성이 점점 중요해지고 있습니다. 그러나 현재의 평가 파이프라인은 여전히 매우 분산되어 있고 밀접하게 연결되어 있어 재현성을 저해하고 불필요한 엔지니어링을 초래합니다. 이러한 문제를 해결하기 위해, 우리는 LLM 기반 에이전트의 평가를 위한 개방형, 경량화되고 확장 가능한 인프라인 AgentCompass를 소개합니다. AgentCompass는 벤치마크, 실행 환경, 그리고 시스템 구성 요소라는 세 가지 독립적인 구성 요소를 중심으로 평가 프로세스를 구성하여, 복잡한 실행 로직을 재구현하지 않고도 유연한 구성을 가능하게 합니다. 또한, AgentCompass는 오류 허용 기능을 갖춘 비동기 런타임과 포상 회피와 같은 미묘한 실패 모드를 투명하게 진단할 수 있는 종합적인 경로 분석 도구를 제공합니다. 20개 이상의 벤치마크를 다섯 가지 능력 차원에서 기본적으로 지원하는 AgentCompass는 연구 커뮤니티에 확장 가능하고 재현 가능한 인프라를 제공하여 에이전트 연구 발전에 기여합니다.
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.