2605.28556v1 May 27, 2026 cs.AI

TASTE: 에이전트 벤치마크의 커버리지 및 난이도 향상을 위한 연구

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Nitay Calderon
Nitay Calderon
Technion
Citations: 378
h-index: 10
Asaf Yehudai
Asaf Yehudai
Citations: 407
h-index: 9
Yotam Perlitz
Yotam Perlitz
Citations: 343
h-index: 9
Tomer Keren
Tomer Keren
Citations: 63
h-index: 3
Michal Shmueli-Scheuer
Michal Shmueli-Scheuer
Citations: 1,031
h-index: 17
Roi Reichert
Roi Reichert
Citations: 0
h-index: 0

에이전트의 기능이 발전함에 따라, 기존의 벤치마크 (예: τ²-Bench)는 점점 포화 상태가 되고 있습니다. 그러나 새로운 벤치마크 작업을 구축하는 것은 여전히 복잡하고 비용이 많이 들며 노동 집약적입니다. 더욱이, 일반적으로 자연어로 작성된 시나리오를 도구 시퀀스로 변환하는 표준적인 접근 방식은 에이전트가 사용하는 도구 사용 패턴의 제한된 부분만 포착합니다. 본 논문에서는 이러한 문제점을 해결하기 위해 작업 생성 프로세스를 역으로 진행합니다. 우리는 TASTE (Task Synthesis from Tool Sequence Evolution)라는 자동화된 방법을 제안하는데, 이는 광범위한 도구 사용 범위를 갖춘 도전적인 작업을 생성합니다. TASTE는 LLM에 의해 판단된 유효성 신호로 학습된 적응형 대비 n-gram 모델을 활용하여, 다양한 도구 조합을 포함하는 유효한 도구 시퀀스를 샘플링할 수 있습니다. 그런 다음 TASTE는 클러스터링을 통해 생성된 풀에서 대표적인 시퀀스를 선택하고, 이를 완전한 벤치마크 작업으로 구현하며, 반복적인 난이도 진화를 통해 개선합니다. TASTE를 사용하여 τ²-Bench의 세 가지 영역을 확장한 도전적인 벤치마크인 τᶜ-Bench를 구축했습니다. 우리는 11개의 에이전트/사용자 LLM 쌍을 평가했으며, τ²-Bench에 거의 포화된 모델들이 저희가 생성한 작업에서 심각한 성능 저하를 보이는 것을 확인했습니다 (예: Gemini-3-Flash는 0.82~0.94에서 0.28~0.61로 하락). 저희가 생성한 작업은 난이도 증가 외에도 에이전트가 실행해야 하는 고유한 도구 조합의 수를 두 배 이상 늘립니다. 이러한 결과는 기존 벤치마크에서의 높은 점수가 견고한 문제 해결 능력보다는 포화 상태를 반영하는 경우가 많다는 것을 시사합니다. TASTE는 어려운, 고도 커버리지를 갖는 벤치마크 작업을 자동으로 생성함으로써, 향후 에이전트에 대한 지속적이고 확장 가능한 평가를 가능하게 합니다.

Original Abstract

As agent capabilities advance, existing benchmarks, such as $τ^2$-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and labor-intensive. Moreover, the standard approach, in which scenarios are first written in natural language and then mapped to tool sequences, captures only a narrow subset of the tool-use patterns agents exercise. In this paper, we address these problems by reversing the task construction process. We propose TASTE: Task Synthesis from Tool Sequence Evolution, an automatic method that generates challenging tasks with broader tool-use coverage. TASTE utilizes an Adaptive Contrastive $n$-gram model trained on LLM-judged validity signals. This enables sampling valid tool sequences that cover a vast range of tool combinations. TASTE then selects representative sequences from the pool via clustering, instantiates them into complete benchmark tasks, and refines them through iterative difficulty evolution. Using TASTE, we construct $τ^c$-Bench, a challenging extension of the three domains of $τ^2$-Bench. We evaluate $11$ agent/user LLM pairs and find that models nearly saturating $τ^2$-Bench suffer severe performance drops on our tasks (e.g., Gemini-3-Flash falls from $0.82\!-\!0.94$ to $0.28\!-\!0.61$). Beyond increasing difficulty, our generated tasks more than double the number of unique tool combinations agents must execute. Our results suggest high scores on existing benchmarks often reflect saturation rather than robust task-solving ability. By automating the generation of difficult, high-coverage benchmarks, TASTE enables continuous, scalable evaluation of future agents.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!