2605.27898v1 May 27, 2026 cs.AI

LLM 에이전트 기능 평가를 위한 통합 프레임워크

A Unified Framework for the Evaluation of LLM Agentic Capabilities

Li Sun
Li Sun
Citations: 10
h-index: 2
Jingyi Yang
Jingyi Yang
Citations: 159
h-index: 6
Pengyu Zhu
Pengyu Zhu
Citations: 35
h-index: 3
S. Su
S. Su
Citations: 216
h-index: 7
Yaxing Lyu
Yaxing Lyu
Citations: 7
h-index: 1
Jing Shao
Jing Shao
Citations: 49
h-index: 3
Lijun Li
Lijun Li
Citations: 1,667
h-index: 10
Yi Liu
Yi Liu
Citations: 17
h-index: 2
Tingfeng Hui
Tingfeng Hui
Citations: 21
h-index: 3
Qi Luo
Qi Luo
Citations: 450
h-index: 9
Xin Yuan
Xin Yuan
Citations: 2
h-index: 1

최근 LLM이 에이전트로 점점 더 많이 활용됨에 따라, 이들의 에이전트 기능을 신뢰성 있게 평가하는 것이 중요해졌습니다. 그러나 보고되는 벤치마크 점수는 종종 모델의 능력과 각 벤치마크가 포함된 구현 방식 모두를 반영하여, 벤치마크 간 결과를 모델 자체의 특성을 나타내는 순수한 측정값으로 해석하기 어렵게 만듭니다. 본 연구에서는 LLM 에이전트 기능을 공정하게 평가할 수 있는 통합 프레임워크를 제시합니다. 이 프레임워크는 통일된 구성 시스템을 기반으로 다양한 벤치마크를 표준화된 지시-도구-환경 형식으로 통합하고, 고정된 ReAct 스타일 아키텍처 내에서 제어 가능한 격리 환경에서 에이전트를 실행하며, 불안정한 실제 환경 대신 선별된 스냅샷을 사용하는 선택적 오프라인 설정을 제공하여 프레임워크 효과와 환경 효과를 별도로 분석할 수 있습니다. 이를 바탕으로 각 벤치마크의 원래 작업 성공 기준을 유지하면서도, 리소스 소비에 대한 통일된 지표와 의사 결정 및 실행 단계에서의 실패 원인 분류 체계를 도입했습니다. 이 프레임워크 내에서, 단일 에이전트, 다중 에이전트, 그리고 안전 관련 시나리오를 아우르는 24개 도메인의 7가지 널리 사용되는 벤치마크를 적용하고, 15개의 모델에 대해 40만 번의 실행과 50억 개의 토큰을 사용하여 대규모 실험 분석을 수행했습니다. 결과는 스캐폴드 선택과 환경 변동성이 벤치마크 결과에 상당한 영향을 미친다는 것을 보여주며, 따라서 우리의 프레임워크가 LLM 자체의 고유한 능력을 프레임워크 및 환경으로 인한 요소로부터 분리할 수 있음을 입증합니다. 또한, 이 프레임워크는 안전 관련 분야를 위한 보안 테스트 환경으로서의 확장 가능성을 보여줍니다. 코드와 벤치마크는 다음 주소에서 이용 가능합니다: https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities, https://huggingface.co/AgentFramework/Unified_Farmework.

Original Abstract

As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model capability and the implementation choices each benchmark is packaged with, making cross-benchmark results difficult to interpret as clean measurements of the underlying model. In this work, we present a unified framework for the fair evaluation of LLM agentic capabilities. Driven by a unified configuration system, the framework integrates diverse benchmarks into a standardized instruction--tool--environment format, executes agents through a fixed ReAct-style architecture within a controllable sandbox, and provides an optional offline setting that replaces volatile live environments with curated snapshots, so that framework effects and environment effects can be analyzed separately. Building on this, we unify the evaluation methodology under each benchmark's original task-success criteria, while introducing unified metrics for resource consumption and a taxonomy for decision- and execution-level failure attribution. Within this framework, we adapt 7 widely used benchmarks spanning 24 domains across single-agent, multi-agent, and safety-critical scenarios, and conduct a large-scale empirical analysis over 400K rollouts and 5B tokens on 15 models. The results show that scaffold choice and environmental volatility materially shift benchmark outcomes in both directions, allowing our framework to disentangle intrinsic LLM capabilities from framework- and environment-induced artifacts. We further demonstrate its extensibility as a secure testbed for safety-critical domains. Codes and benchmarks at are available at https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities, https://huggingface.co/AgentFramework/Unified_Farmework.

3 Citations
0 Influential
25 Altmetric
13.9 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!