정적인 순위표를 넘어서: LLM 에이전트 평가를 위한 예측 타당성
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
에이전트 벤치마크는 빠르게 증가하고 있지만, 단일 벤치마크가 실제 배포 환경에서 드러나는 다양한 측면 중 네다섯 개 이상을 포괄하는 경우는 드뭅니다. 본 논문에서는 현재까지 가장 광범위하고 심층적인 MCP 기반 산업용 에이전트 벤치마크 연구 결과를 종합적으로 분석합니다. 이 연구는 새로운 자산 클래스(멀티모달 시각 정보 포함), 다양한 오케스트레이션 방식, 검색 전략, 추론 모드, 인프라 최적화 및 평가 방법론에 대한 14개의 병렬 구현 연구를 다룹니다. 이러한 연구와 함께 기존의 7가지 에이전트 벤치마크 결과를 종합하여 분석한 결과, 집계 점수를 기반으로 한 순위표는 실제 배포 환경에서 에이전트를 평가하는 데 있어 중요한 정보를 누락시킬 수 있음을 주장합니다. 집계 점수에 따른 순위는 일반화되지 않은 환경에서도 그대로 유지되지 않으며, 최근 공개된 경쟁 대회 결과를 분석한 결과 이러한 순위 불안정성에 대한 직접적인 경험적 증거를 제시합니다. 본 연구에서는 예측 타당성, 즉 학습 데이터에서의 순위와 실제 데이터에서의 순위 간의 상관관계를 기준으로 에이전트 평가 시스템을 구성하는 방법을 제안하고, HELM 및 후속 모델에서 다루는 배포 관련 주요 측면들을 측정할 수 있는 12단계 측정 도구를 제시합니다. 이 방법론은 세 가지 검증 가능한 일반화 기준과 명확한 임계값을 통해 구현되며, 현재까지의 증거는 이를 부분적으로 뒷받침하지만, 아직 충분히 확정하기에는 부족합니다. 마지막으로, 본 논문에서는 사전 등록된 파일럿 연구 설계와 향후 에이전트 벤치마크가 제공해야 할 내용에 대한 비전을 제시합니다.
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.