2606.11070v1 Jun 09, 2026 cs.CL

T1-벤치: 실제 환경에서의 다중 시나리오 에이전트 성능 측정

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shi-Xiong Zhang
Shi-Xiong Zhang
Citations: 122
h-index: 5
Sambit Sahu
Sambit Sahu
Citations: 87
h-index: 5
G. Winata
G. Winata
Citations: 74
h-index: 4
Anirban Das
Anirban Das
Citations: 123
h-index: 5
Paresh Dashore
Paresh Dashore
Citations: 11
h-index: 2
A. Chakraborty
A. Chakraborty
Citations: 53
h-index: 2
Yuzhen Lin
Yuzhen Lin
Citations: 62
h-index: 5
Swasthi P. Rao
Swasthi P. Rao
Citations: 4
h-index: 1
Houhan Lu
Houhan Lu
Citations: 0
h-index: 0
Nadia Bathaee
Nadia Bathaee
Citations: 22
h-index: 3
Sriharsha Hatwar
Sriharsha Hatwar
Citations: 2,713
h-index: 2
Anmol Jain
Anmol Jain
Citations: 0
h-index: 0
Kshitij Tayal
Kshitij Tayal
Citations: 318
h-index: 11
Xiuzhu Lin
Xiuzhu Lin
Citations: 9
h-index: 1

최근 대규모 언어 모델(LLM)의 추론 및 도구 활용 능력 발전으로 더욱 강력한 에이전트 시스템 개발이 가능해졌습니다. 그러나 기존 벤치마크는 작업 복잡성, 현실성 및 도메인 다양성 측면에서 한계가 있으며, 여러 도메인을 포괄하는 상호 작용을 제대로 반영하지 못하여 지속적인 추론과 조율이 필요한 실제 다단계 환경에서의 에이전트 성능 평가에 어려움이 있습니다. 이러한 문제점을 해결하기 위해, 우리는 T1-벤치를 소개합니다. T1-벤치는 실제 고객 응대 및 다중 도메인 환경에서 에이전트 시스템을 평가하기 위한 고정밀, 종합적인 벤치마크입니다. 이 벤치마크는 체계적인 추론을 요구하는 다단계 사용자-어시스턴트 상호 작용을 포함하며, 총 25개의 다양한 난이도 도메인에 걸쳐 작업 복잡성과 평가의 엄격함을 크게 향상시켰습니다. 우리는 12개의 독점 모델 및 오픈 가중치 모델을 사용하여 T1-벤치를 평가하고, 복잡한 다단계 환경에서 에이전트의 행동, 도구 활용 및 대화 품질을 평가하기 위한 재현 가능하고 표준화된 프레임워크를 제공합니다. 또한 자동 평가와 함께 인간 판단을 추가하여 정성적 성능 평가를 강화했습니다. 전반적으로 T1-벤치는 시뮬레이션된 다중 도메인 환경에서 작업 복잡성, 상호 작용 깊이 및 도메인 커버리지를 증가시켜 기존 벤치마크를 크게 발전시켰습니다. 에이전트 시스템에 대한 향후 연구를 지원하기 위해, 데이터와 평가 코드를 오픈 소스로 공개할 예정입니다.

Original Abstract

Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain limited in task complexity, realism, and domain diversity, and often fail to capture interactions that span multiple domains, limiting their ability to evaluate agents in realistic multi-step settings that require sustained reasoning and coordination. To address these limitations, we introduce T1-Bench, a high-fidelity, comprehensive benchmark for evaluating agentic systems in realistic customer-facing, multi-domain environments, featuring interleaved scenarios that require structured reasoning across multi-turn user-assistant interactions and substantially increasing both compositional complexity and evaluative rigor across 25 domains of varying difficulty. We evaluate T1-Bench using 12 proprietary and open-weight models, providing a reproducible and standardized framework for assessing agent behavior, tool utilization, and conversational quality in complex, multi-step environments. We further complement automatic evaluation with human judgments to strengthen the assessment of qualitative performance. Overall, T1-Bench substantially advances prior benchmarks by increasing task complexity, interaction depth, and domain coverage in simulated multi-domain environments. To facilitate future research on agentic systems, we will publicly release data and evaluation code as open source.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!