GABench: 그래프 분석 작업에서 LLM 에이전트 평가를 위한 종합적인 벤치마크
GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks
대규모 언어 모델(LLM) 에이전트는 계획, 도구 사용 및 외부 환경과의 상호 작용 능력이 점점 더 발전하고 있습니다. 이러한 에이전트는 일반적으로 상태를 관리하고 다단계 실행을 조정하는 데 사용되는 프레임워크의 지원을 받습니다. 그래프 분석은 에이전트가 데이터에 접근하고 그래프 환경에서 작업을 수행해야 하므로, LLM 에이전트의 능력을 평가하기 위한 유망한 환경을 제공합니다. 그러나 기존의 LLM을 위한 그래프 벤치마크는 그래프 작업 및 유형 측면에서 제한적인 범위를 가지며, 이는 LLM 에이전트를 종합적으로 평가하는 데 어려움을 야기합니다. 또한, 이러한 벤치마크들은 일반적으로 그래프 정보를 프롬프트에 직접 제공하는 텍스트 기반 질의 응답 방식으로 그래프 분석을 구성하여, 엔드투엔드 에이전트 능력을 평가하는 데 한계가 있습니다. 이러한 제한 사항을 해결하기 위해, 우리는 에이전트 중심 그래프 분석을 위한 종합적인 벤치마크인 GABench를 소개합니다. GABench는 세 가지 유형의 그래프를 포함하며, 그래프 검색, 그래프 이론, 그래프 머신 러닝 및 그래프 개방형 질의 응답이라는 네 가지 범주의 그래프 분석 작업을 다룹니다. GABench는 또한 그래프 데이터를 접근하고 다양한 그래프 작업을 수행하기 위한 84개의 실행 가능한 도구를 제공합니다. 이러한 도구를 기반으로, 우리는 에이전트 중심 그래프 분석 작업 생성 파이프라인을 개발하고 검증 가능한 정답을 가진 10,400개의 작업을 구축했습니다. GABench를 사용하여 다양한 최첨단 LLM 및 에이전트 프레임워크를 평가했습니다. 우리의 실험 결과는 세 가지 주요 결과를 보여줍니다: (1) 기존의 LLM 에이전트는 여전히 복잡한 그래프 분석 작업에 어려움을 겪습니다. (2) 프레임워크 선택은 성능에 상당한 영향을 미치지만, 기존 프레임워크는 복잡한 그래프 작업에서 여전히 제한적입니다. (3) 그래프 분석은 도구 호출의 품질이 양보다 더 중요합니다. 우리의 연구 결과는 그래프 분석을 위한 LLM 에이전트 개발 및 평가에 대한 실질적인 통찰력을 제공합니다.
Large language model (LLM) agents are increasingly capable of planning, using tools, and interacting with external environments. They are typically supported by harnesses, which manage state and coordinate multi-step execution. Graph analysis provides a promising setting for evaluating their agentic capabilities, because it requires agents to access data and execute operations in a graph environment. However, existing graph benchmarks for LLMs provide limited coverage of graph tasks and graph types, making it difficult to comprehensively evaluate LLM agents. Moreover, they typically formulate graph analysis as text-based question answering, where graph information is directly provided in the prompt, limiting the evaluation of end-to-end agentic capabilities. To address these limitations, we introduce GABench, a comprehensive benchmark for agentic graph analysis. GABench spans three graph types and covers four graph analysis task categories: graph retrieval, graph theory, graph machine learning, and graph open-ended question answering. GABench also provides 84 executable tools for accessing graph data and performing diverse graph operations. Building on these tools, we develop an agentic graph analysis task generation pipeline and construct 10,400 tasks with verifiable ground truth.Using GABench, we evaluate a range of frontier LLMs and agent harnesses. Our experiments reveal three key findings: (1) Existing LLM agents still struggle with complex graph analysis tasks. (2) Harness choice significantly affects performance, yet existing harnesses remain limited on complex graph tasks. (3) Graph analysis depends more on tool-call quality than quantity. Our findings provide practical insights into the development and evaluation of LLM agents for graph analysis.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.