2607.05174v1 Jul 06, 2026 cs.AI

AgentGym2: 비이상화된 실제 환경에서 대규모 언어 모델 에이전트 성능 평가

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

Xuanjing Huang
Xuanjing Huang
Citations: 3,773
h-index: 33
Zhiheng Xi
Zhiheng Xi
Citations: 1,507
h-index: 17
Zhihao Zhang
Zhihao Zhang
Citations: 340
h-index: 7
Qi Zhang
Qi Zhang
Citations: 1,944
h-index: 25
Dongrui Liu
Dongrui Liu
Citations: 49
h-index: 3
Junjie Ye
Junjie Ye
Fudan University
Citations: 1,596
h-index: 17
Tao Gui
Tao Gui
Citations: 450
h-index: 7
Honglin Guo
Honglin Guo
Fudan University
Citations: 554
h-index: 10
Jiaming Ji
Jiaming Ji
Citations: 1,019
h-index: 18
Dingwei Zhu
Dingwei Zhu
Citations: 11
h-index: 2
Jiazheng Zhang
Jiazheng Zhang
Citations: 121
h-index: 6
Xin Guo
Xin Guo
Citations: 2,352
h-index: 6
Jiaqi Liu
Jiaqi Liu
Citations: 93
h-index: 2
Junzhe Wang
Junzhe Wang
Citations: 2,355
h-index: 7
Dingwen Yang
Dingwen Yang
Citations: 175
h-index: 3
Jixuan Huang
Jixuan Huang
Citations: 156
h-index: 4
Baodai Huang
Baodai Huang
Citations: 67
h-index: 2
Tinggang Chen
Tinggang Chen
Citations: 0
h-index: 0
Zhonghang Lu
Zhonghang Lu
Citations: 0
h-index: 0
Chenyu Liu
Chenyu Liu
Citations: 0
h-index: 0
Jiajun Sun
Jiajun Sun
Citations: 24
h-index: 3
Yuming Yang
Yuming Yang
Fudan University
Citations: 234
h-index: 10
Guohao Li
Guohao Li
Citations: 0
h-index: 0
Minghe Gao
Minghe Gao
Citations: 287
h-index: 9

언어 모델 기반 에이전트는 빠르게 발전하고 있으며, 점점 더 많은 생산 환경에 적용되고 있습니다. 이러한 추세는 엄격하고 현실적인 평가의 필요성을 강조합니다. 그러나 대부분의 기존 벤치마크는 단순화된 이상적인 환경에서 에이전트를 평가합니다. 일반적으로 미리 정의된 도구 인터페이스에 의존하며, 중요한 단계를 간과하고 입력 데이터가 깨끗하고 완전하다고 가정합니다. 결과적으로 이러한 벤치마크는 실제 배포 환경의 어려움을 과소평가하는데, 실제 환경에서는 불확실성과 노이즈가 만연하며 에이전트는 새로운 도구를 발견하기 위해 능동적으로 환경을 탐색해야 합니다. 이러한 격차를 해소하기 위해 우리는 AgentGym2라는 새로운 평가 프레임워크를 제시합니다. AgentGym2는 실제 업무 요구 사항에 기반한 작업 인스턴스를 포함하고 있으며, 추론 및 계획 능력 외에도 에이전트가 전체 절차를 실행하고, 탐색을 통해 도구를 발견하며, 새롭지 않은 작업을 위해 도구를 조합하고, 노이즈와 불완전한 정보에 대한 견고성을 유지하는 능력을 측정합니다. 15개의 독점적 및 오픈 소스 모델에 대한 실험 결과는 Gemini 및 GPT-5와 같은 최첨단 시스템조차도 AgentGym2에서 어려움을 겪으며, 이는 현재 에이전트의 능력과 실제 응용 분야의 요구 사항 간에 상당한 격차가 있음을 보여줍니다.

Original Abstract

Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They typically rely on pre-packaged tool interfaces, overlook critical steps, and assume inputs are clean and fully specified. Consequently, they understate the difficulty of real deployments, where uncertainty and noise are ubiquitous and agents must proactively explore the environment to uncover new tools. To bridge this gap, we present AgentGym2, a new evaluation framework with task instances grounded in real-world end-to-end working demands. Beyond reasoning and planning, it measures agents' ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information. Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2, revealing a substantial gap between the capability of current agents and the demands of real-world applications.

1 Citations
0 Influential
16.5 Altmetric
83.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!