OmniaBench: 다양한 시나리오에서의 범용 인공지능 에이전트 성능 평가
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
대규모 언어 모델은 점점 더 텍스트 생성기로부터 사용자 요청을 이해하고, 외부 도구를 활용하며, 상호작용을 통해 복잡한 작업을 수행할 수 있는 범용 에이전트로 진화하고 있습니다. 그러나 기존의 에이전트 벤치마크는 종종 제한적인 시나리오, 도구 생태계 또는 상호 작용 형식에 초점을 맞추어 모델의 다양한 애플리케이션 환경에서의 기능을 체계적으로 평가하기 어렵게 만듭니다. 본 논문에서는 명시적인 상태 공간을 갖춘 다양한 시나리오에서 범용 에이전트를 평가하기 위한 벤치마크인 OmniaBench를 소개합니다. 우리는 앱 스토어, 제품 문서, 산업 자료, 웹 검색 및 인간의 검토를 통해 애플리케이션 지향적 시나리오 지식을 추출하여 ToC (Text-to-Code), ToB (Text-to-Business) 및 ToE (Text-to-Entity)를 포함하는 90개의 레벨-1 도메인과 354개의 레벨-2 도메인을 포괄하는 계층적 분류 체계를 구축했습니다. 이 분류 체계를 기반으로 실행 가능한 환경을 구성하고, DAG, DAG-S, Solver 및 Program의 네 가지 상호 보완적인 방법을 통해 단일 턴 및 다중 턴 작업을 생성합니다. OmniaBench는 또한 10차원의 능력 분류 체계와 8가지의 조합 가능한 기본 난이도 요소를 도입하여 세밀한 평가 및 분석을 지원합니다. 결과적으로 생성된 데이터셋에는 1,431개의 작업이 포함되어 있으며, 평가 비용을 줄이고 공개 후 전체 데이터셋의 잠재적인 오염을 완화하기 위해 설계된 644개의 어려운 작업 하위 집합도 포함됩니다. OmniaBench는 현재 최고 성능 모델에도 상당한 어려움을 제시하며, Claude-Sonnet-5와 GPT-5.6-Sol조차도 Overall Pass@1 점수가 각각 58.54점과 57.14점에 불과합니다. 추가적인 분석 결과, 도메인 및 능력 간에 명확한 차이가 있으며, 계획 수립, 제약 조건 유지 및 적응적 수정에 있어 지속적인 한계가 있음을 보여줍니다. OmniaBench는 범용 에이전트의 기능 경계를 특성화하기 위한 포괄적이고 진단적인 벤치마크를 제공합니다.
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across diverse scenarios with explicit state spaces. We derive application-oriented scenario knowledge from app stores, product documents, industry resources, Web retrieval, and human refinement, forming a hierarchical taxonomy that spans ToC, ToB and ToE with 90 level-1 and 354 level-2 domains. Based on this taxonomy, we construct executable environments and synthesize single-turn and multi-turn tasks through four complementary routes: DAG, DAG-S, Solver, and Program. OmniaBench further introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis. The resulting dataset contains 1,431 tasks, together with a challenging subset of 644 tasks designed to reduce evaluation cost and mitigate potential contamination of the full set after public release. The bench presents substantial challenges to current frontier models, with even Claude-Sonnet-5 and GPT-5.6-Sol achieving Overall Pass@1 scores of only 58.54 and 57.14, respectively. Further analyses reveal clear differences across domains and capabilities, as well as persistent limitations in planning, constraint maintenance, and adaptive correction. OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.