AgentBeats: 개방성, 표준화 및 재현성을 위한 에이전트 평가의 에이전트 기반 접근 방식
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
에이전트 시스템은 다양한 분야에서 빠르게 발전하고 있지만, 그 평가는 여전히 단편적입니다. 대부분의 벤치마크는 고정된 LLM 중심 프레임워크에 의존하며, 이는 복잡한 통합을 요구하고, 테스트 환경과 실제 환경 간의 불일치를 야기하며, 다양한 에이전트 설계 간의 공정한 비교를 제한합니다. 근본적인 문제는 개방적이고, 에이전트 독립적인 평가 인터페이스가 부족하다는 점입니다. 우리는 에이전트 기반 에이전트 평가(AAA)를 제안하는데, 여기서 평가는 심판 에이전트에 의해 수행되며, 모든 참가자는 표준화된 프로토콜(A2A: 작업 관리 및 MCP: 도구 접근)을 통해 상호 작용합니다. 기존 벤치마킹은 벤치마크 인터페이스와 에이전트 인터페이스라는 두 개의 별도 인터페이스를 정의하는 반면, AAA는 단일 인터페이스만 필요합니다. 이를 통해 평가 로직과 에이전트 구현이 분리된 범용적이고 통합적인 프레임워크가 구축되어, 재현 가능하고 상호 운용 가능한 다중 에이전트 평가를 가능하게 합니다. 또한, 우리는 AAA의 구체적인 구현체인 AgentBeats를 소개합니다. AgentBeats는 개방성, 개인 정보 보호 및 재현성에 대한 실제 제약 조건을 고려하여 표준화된 평가가 가능하도록 하는 5가지 실용적인 운영 모드를 제시합니다. 우리의 설계를 대규모로 평가하기 위해 두 가지 연구를 수행했습니다. 첫 번째는 298개의 심판 에이전트와 12개 범주, 그리고 독립적인 참가자로부터 온 467개의 대상 에이전트를 포함하는 5개월간의 공개 경쟁으로, AAA가 다양한 벤치마크에 적용될 수 있음을 보여줍니다. 두 번째는 코딩 에이전트에 대한 사례 연구로, 에이전트 기반 평가가 기존의 공개 기록과 일관성을 유지하면서 이전에 발견되지 않았던 직접적인 비교 결과를 제공하여 에이전트 설계에 대한 연구적 통찰력을 얻을 수 있음을 확인합니다. 커뮤니티 규모의 현장 연구와 제어된 코딩 사례 연구를 결합하여, AAA가 다양한 시나리오에서 대규모로 적용 가능하며, 실용성과 정확성을 모두 갖춤을 검증했습니다. AAA와 AgentBeats는 개방적이고 표준화된, 그리고 재현 가능한 에이전트 평가를 위한 명확한 경로를 제시합니다.
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.