MCPEvol-Bench: MCP 서버의 동적 변화에 따른 LLM 에이전트 성능 벤치마킹
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
모델 컨텍스트 프로토콜(MCP) 서버가 LLM과 외부 도구를 연결하는 핵심 인프라로 부상함에 따라, 기존 벤치마크는 실제 MCP 서버를 활용하여 LLM 에이전트의 도구 사용 능력을 평가합니다. 그러나 이러한 벤치마크는 MCP 서버 내에서 지속적으로 변화하는 도구 인터페이스 및 기능성을 간과하여, 에이전트가 변화하는 도구 환경에 얼마나 잘 적응하는지를 제대로 반영하지 못하는 부정확한 평가를 초래합니다. 이러한 문제점을 해결하기 위해, 우리는 LLM 에이전트의 문제 해결 능력을 동적으로 변화하는 도구 세트를 기반으로 평가하는 새로운 벤치마크인 **MCPEvol-Bench**를 소개합니다. 대규모 실증 연구에서 영감을 받아, 우리는 123개의 MCP 서버 내에서 현실적인 도구 진화를 시뮬레이션하기 위해 11가지 변이 연산자를 제안했습니다. 우리는 최첨단 LLM 12개를 다양한 버전의 MCP 서버에서 벤치마킹한 결과, 심지어 최고 수준의 모델조차도 변화하는 도구에 적응하는 데 어려움을 겪는다는 것을 확인했습니다. 예를 들어, GPT-5.4와 Claude-Sonnet-4-6는 진화된 MCP 서버에서 각각 13.7% 및 14.4%의 성능 저하를 보였으며, 이는 계획 및 추론 오류의 상당한 증가를 동반했습니다. 이러한 결과는 LLM 기반 워크플로우의 취약점을 강조하며, MCPEvol-Bench가 동적 도구 환경에서 에이전트의 적응성을 평가하는 표준으로 자리매김할 수 있음을 시사합니다.
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's adaptability in changing tool landscapes. To bridge this gap, we introduce \textbf{MCPEvol-Bench}, a novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution. Inspired by large-scale empirical study, we propose 11 mutation operators to simulate realistic tool evolution within 123 MCP servers. We benchmark 12 state-of-the-art LLMs on multiple versions of MCP servers, revealing that even frontier models struggle to adapt to evolving tools. For instance, GPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7\% and 14.4\% in evolved MCP servers, respectively, accompanied by substantial increases in planning and reasoning errors. These findings highlight the vulnerability of LLM-driven workflows, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.