Messier: 교차 벤치마크 에이전트 평가를 위한 고해상도 데이터셋
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
인터랙티브 환경에서 AI 에이전트를 평가하는 것은 단편적인 작업, 안내 요소, 검증 도구 및 점수 규칙으로 인해 어려움을 겪습니다. 기존 연구는 주로 제한된 환경에 초점을 맞추거나, 규모가 작거나, 비용이 많이 드는 재실행을 필요로 하며, 이는 대부분의 실증적 결과를 비교 불가능하게 만듭니다. 본 논문에서는 30개의 벤치마크, 714개의 에이전트, 11,891개의 작업 및 74,205개의 검증 도구를 포함하는 총 957,253개의 레코드를 담은 통합 데이터셋인 Messier를 소개합니다. Messier는 공개된 벤치마크 점수를 통합하고, 최근 법률 벤치마크를 포함하여 대표성이 낮은 6개 전문 및 과학 분야에 걸쳐 5개의 에이전트를 사용하여 추가적인 실험 결과를 제공합니다. 각 레코드는 모델, 안내 요소, 환경, 작업, 검증 도구 및 집계 규칙을 기준으로 표준화되며, 직업 및 산업 분석을 위한 SOC/NAICS 분류를 포함합니다. 본 데이터셋을 사용하여 벤치마크 유형에 따른 발전 수준이 균일하지 않음을 보여줍니다. "함수 호출"은 포화 상태이며, "프로그래밍"은 가장 빠르게 개선되고 있으며, "기업 워크플로우"는 여전히 가장 어려운 것으로 나타났습니다. 또한, 반사실적 재점수를 통한 분석 결과, 다중 검증 도구를 사용하는 작업에서 엄격한 모든 통과 점수 집계 방식은 발전을 가리고 에이전트 순위를 인위적으로 변경할 수 있음을 확인했습니다. 이러한 표준화된 레코드를 기반으로 Epoch의 평가 능력 지수 순위에 Spearman 상관 계수 0.81로 일치하는 기능 수준을 도출했으며, 이를 특정 분야, 직업, 액션 공간 또는 검증 도구 유형에 따라 세분화할 수 있습니다. Messier는 에이전트 기능 수준 측정, 벤치마크 감사 및 평가 실패에 대한 정밀 분석을 위한 기본적인 재사용 가능한 인프라를 제공합니다.
Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six underrepresented professional and scientific domains, including a recent legal benchmark. Each record is standardized by model, scaffold, environment, task, verifier, and aggregation rule, with SOC/NAICS classifications for occupational and industry analysis. Using this corpus, we show frontier progress is uneven across benchmark types, with "function calling" saturated, "programming" improving the fastest, and "enterprise workflows" remaining the most challenging. Furthermore, counterfactual rescoring shows that strict all-pass aggregation in multi-verifier tasks can obscure progress and artificially alter agent rankings. From these standardized records, we derive capability scales that align with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.81 and can be specialized by domain, occupation, action space, or verifier type. Messier provides a foundational, reusable infrastructure for agent capability scaling, benchmark auditing, and fine-grained analysis of evaluation failures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.