AgentGym2: 비이상화된 실제 환경에서 대규모 언어 모델 에이전트 성능 평가
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
언어 모델 기반 에이전트는 빠르게 발전하고 있으며, 점점 더 많은 생산 환경에 적용되고 있습니다. 이러한 추세는 엄격하고 현실적인 평가의 필요성을 강조합니다. 그러나 대부분의 기존 벤치마크는 단순화된 이상적인 환경에서 에이전트를 평가합니다. 일반적으로 미리 정의된 도구 인터페이스에 의존하며, 중요한 단계를 간과하고 입력 데이터가 깨끗하고 완전하다고 가정합니다. 결과적으로 이러한 벤치마크는 실제 배포 환경의 어려움을 과소평가하는데, 실제 환경에서는 불확실성과 노이즈가 만연하며 에이전트는 새로운 도구를 발견하기 위해 능동적으로 환경을 탐색해야 합니다. 이러한 격차를 해소하기 위해 우리는 AgentGym2라는 새로운 평가 프레임워크를 제시합니다. AgentGym2는 실제 업무 요구 사항에 기반한 작업 인스턴스를 포함하고 있으며, 추론 및 계획 능력 외에도 에이전트가 전체 절차를 실행하고, 탐색을 통해 도구를 발견하며, 새롭지 않은 작업을 위해 도구를 조합하고, 노이즈와 불완전한 정보에 대한 견고성을 유지하는 능력을 측정합니다. 15개의 독점적 및 오픈 소스 모델에 대한 실험 결과는 Gemini 및 GPT-5와 같은 최첨단 시스템조차도 AgentGym2에서 어려움을 겪으며, 이는 현재 에이전트의 능력과 실제 응용 분야의 요구 사항 간에 상당한 격차가 있음을 보여줍니다.
Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They typically rely on pre-packaged tool interfaces, overlook critical steps, and assume inputs are clean and fully specified. Consequently, they understate the difficulty of real deployments, where uncertainty and noise are ubiquitous and agents must proactively explore the environment to uncover new tools. To bridge this gap, we present AgentGym2, a new evaluation framework with task instances grounded in real-world end-to-end working demands. Beyond reasoning and planning, it measures agents' ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information. Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2, revealing a substantial gap between the capability of current agents and the demands of real-world applications.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.