AgentPulse: 배포 환경에서 AI 에이전트를 평가하기 위한 지속적인 다중 신호 프레임워크
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
정적 벤치마크는 AI 에이전트가 특정 시점에 수행할 수 있는 작업을 측정하지만, 실제 배포 환경에서의 채택, 유지 관리 및 사용 경험은 측정하지 못합니다. 본 연구에서는 AgentPulse를 소개합니다. AgentPulse는 GitHub, 패키지 저장소, IDE 마켓플레이스, 소셜 플랫폼, 벤치마크 리더보드 등 18가지 실시간 신호를 종합하여 10가지 작업 부문에서 50개의 에이전트를 4가지 요소(벤치마크 성능, 채택 신호, 커뮤니티 여론, 생태계 건강)를 기준으로 평가하는 지속적인 평가 프레임워크입니다. 본 연구에서는 세 가지 분석을 통해 프레임워크의 유효성을 검증했습니다. 네 가지 요소는 대체적으로 상호 보완적인 정보를 제공합니다 (n=50; 채택-생태계 간의 상관관계가 최대 0.61이며, 그 외 모든 경우 |ρ| ≤ 0.37). 순환성을 제어한 테스트(n=35) 결과, GitHub에서 얻은 신호가 포함되지 않은 벤치마크+여론 하위 요소가 GitHub 스타 수(ρs=0.52, p<0.01), Stack Overflow 질문 건수(ρs=0.49, p<0.01)와 같은 외부 채택 지표를 예측할 수 있으며, VS Code 설치 건수(ρs=0.44, p<0.05)는 예시로 제시되었으며, 35개 에이전트 중 11개 에이전트만 설치 건수가 0이 아닌 경우입니다. 또한, SWE-bench 점수가 공개된 11개 에이전트 그룹에서는 종합 점수와 벤치마크 점수 간의 상관관계가 거의 나타나지 않았습니다 (ρs=0.25; 11개 에이전트 중 9개가 최소 2단계 이상 순위 변동). 이는 폐쇄형 소스 고성능 에이전트 그룹 내에서 채택률과 성능 간에 강한 부정적 상관관계가 있기 때문입니다. 따라서 프레임워크의 유효성 검증은 SWE-bench와 관련된 부분보다 더 넓은 범위의 n=35 그룹을 통해 이루어졌습니다. AgentPulse는 벤치마크에서 확인되지 않는 배포 관련 신호를 제공하며, 이는 절대적인 순위를 나타내는 것이 아니라 평가 방법론입니다. 본 프레임워크, 수집된 모든 신호, 점수 결과, 평가 도구는 CC BY 4.0 라이선스 하에 공개됩니다.
Static benchmarks measure what AI agents can do at a fixed point in time but not how they are adopted, maintained, or experienced in deployment. We introduce AgentPulse, a continuous evaluation framework scoring 50 agents across 10 workload categories along four factors (Benchmark Performance, Adoption Signals, Community Sentiment, and Ecosystem Health) aggregated from 18 real-time signals across GitHub, package registries, IDE marketplaces, social platforms, and benchmark leaderboards. Three analyses ground the framework. The four factors capture largely complementary information (n=50; $ρ_{\max}=0.61$ for Adoption-Ecosystem, all others $|ρ| \leq 0.37$). A circularity-controlled test (n=35) shows the Benchmark+Sentiment sub-composite, which contains no GitHub-derived signals, predicts external adoption proxies it does not aggregate: GitHub stars ($ρ_s=0.52$, $p<0.01$) and Stack Overflow question volume ($ρ_s=0.49$, $p<0.01$), with VS Code installs ($ρ_s=0.44$, $p<0.05$) reported as illustrative given that only 11 of 35 agents have non-zero installs. On the n=11 subset with published SWE-bench scores, composite and benchmark-only rankings are nearly uncorrelated ($ρ_s=0.25$; 9 of 11 agents shift by at least 2 ranks), driven by a strong negative Adoption-Capability correlation among closed-source high-capability agents within this subset. This is precisely why we rest the framework's validity claim on the broader n=35 test rather than the SWE-bench overlap. AgentPulse surfaces deployment signal absent from benchmarks; it is a methodology, not a ground-truth ranking. The framework, all collected signals, scoring outputs, and evaluation harness are released under CC BY 4.0.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.