2608.12282v1 Aug 12, 2026 cs.AI

VAKRA: API 및 정보 검색 에이전트의 다중 단계 추론 능력을 평가하기 위한 벤치마크

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Anupama Murthi
Anupama Murthi
Citations: 135
h-index: 4
Abhinav Jain
Abhinav Jain
Citations: 28
h-index: 4
Ankita Rajaram Naik
Ankita Rajaram Naik
University of Massachusetts Amherst
Citations: 568
h-index: 5
Benjamin Elder
Benjamin Elder
Citations: 52
h-index: 5
Siyu Huo
Siyu Huo
Citations: 40
h-index: 4
Praveen Venkateswaran
Praveen Venkateswaran
Citations: 98
h-index: 6
Abdulhamid A. Adebayo
Abdulhamid A. Adebayo
Citations: 117
h-index: 6
Danish Contractor
Danish Contractor
Citations: 85
h-index: 4

기업 환경에 배포되는 에이전트는 구조화된 API와 문서 컬렉션을 활용하여 추론해야 하지만, 기존의 벤치마크는 이러한 기능을 개별적으로 평가합니다. 본 논문에서는 VAKRA (e extbf{V}aluating extbf{A}PI and extbf{K}nowledge extbf{R}etrieval extbf{A}gents)라는 벤치마크를 소개합니다. 이 벤치마크는 62개의 도메인에 걸쳐 8,000개 이상의 실행 가능한 API를 포함하며, 난이도가 점진적으로 증가하는 세 가지 유형의 작업을 포함합니다: 다양한 API 상호 작용 방식, 구조화된 API를 통한 다중 단계 추론, 그리고 자연어 기반의 정책 제약 조건 하에서 여러 소스를 활용한 추론입니다. 예측된 도구 호출은 실제 API를 통해 재실행하여 정확성을 검증하며, 여러 개의 유효 경로를 수용합니다. 모델의 기능과 에이전트 아키텍처를 분리하기 위해 고정된 ReAct 프레임워크를 사용하여 최첨단 및 공개 가중치 모델을 평가한 결과, 가장 성능이 좋은 모델조차도 단일 단계의 엔드포인트 스타일 작업에서 70.4%의 정확도를 보였으며, 복합 API에서는 50~51%로 떨어지는 것을 확인했습니다. 추론 깊이가 증가함에 따라 성능은 50% 이상 저하되며, 정책 제약 조건이 있는 질문에서는 심각한 오류가 발생합니다 (예: 답변 불가능한 질문에서 2.4%). 분석 결과, 대부분의 오류는 도구 호출 메커니즘보다는 언어를 통한 추론 과정, 즉 개체 식별 및 여러 소스 간의 연관성 설정과 관련된 부분에서 발생하는 것으로 나타났습니다. 코드와 데이터셋은 각각 https://github.com/IBM/VAKRA 및 https://huggingface.co/datasets/ibm-research/VAKRA 에서 확인할 수 있습니다.

Original Abstract

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!