2603.28569v1 Mar 30, 2026 cs.LG

CirrusBench: 실제 클라우드 서비스 환경에서 LLM 기반 에이전트의 정확성 평가를 넘어선 연구

CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yipeng Yu
Yipeng Yu
Citations: 8
h-index: 2
Guang Hu
Guang Hu
Citations: 33
h-index: 4
Chen Shen
Chen Shen
Citations: 3
h-index: 1
Xing-Ying Liu
Xing-Ying Liu
Citations: 1
h-index: 1
Jing Gu
Jing Gu
Citations: 7
h-index: 2
Hangyi Sun
Hangyi Sun
Citations: 1
h-index: 1
Jianfeng Liu
Jianfeng Liu
Citations: 132
h-index: 3
Mingyue Pu
Mingyue Pu
Citations: 41
h-index: 3
Yu Wang
Yu Wang
Citations: 8
h-index: 2
Z. Xiao
Z. Xiao
Citations: 17
h-index: 2
Rui Xie
Rui Xie
Citations: 57
h-index: 5
Longjiu Luo
Longjiu Luo
Citations: 1
h-index: 1
Qianrong Wang
Qianrong Wang
Citations: 182
h-index: 5
Gurong Cui
Gurong Cui
Citations: 1
h-index: 1
Hongli Qiao
Hongli Qiao
Citations: 26
h-index: 2
Wenlian Lu
Wenlian Lu
Citations: 46
h-index: 2
Weiting Liu
Weiting Liu
Citations: 78
h-index: 5

대규모 언어 모델(LLM)의 에이전트 기능이 향상됨에 따라, 클라우드 서비스와 같이 높은 수준의 기술적 복잡성과 장기적인 의존성을 갖는 실제 응용 분야에 LLM 기반 에이전트가 활용되고 있습니다. 이러한 환경에서는 고객 만족도를 높이기 위해 안정성과 문제 해결 효율성이 매우 중요합니다. 그러나 기존의 LLM 기반 에이전트 벤치마크는 실제 고객의 다양한 입력과 예측 불가능성을 제대로 반영하지 못하는 인공 환경에 주로 의존하며, 실제 환경에 적용하기 위해 필수적인 문제 해결 효율성을 간과하는 경우가 많습니다. 이러한 격차를 해소하기 위해, 실제 클라우드 서비스 티켓에서 수집한 데이터를 기반으로 하는 새로운 평가 프레임워크인 CirrusBench를 소개합니다. CirrusBench는 기술 서비스 환경에 내재된 복잡한 다중 단계 논리 흐름과 현실적인 도구 의존성을 그대로 유지합니다. 본 연구는 실행 정확성뿐만 아니라, 정규화 효율성 지수(Normalized Efficiency Index) 및 다중 턴 지연 시간(Multi-Turn Latency)과 같은 고객 중심 지표를 도입하여 에이전트의 성공 여부를 정의하고, 서비스 품질을 정량적으로 측정합니다. CirrusBench 프레임워크를 사용한 실험 결과, 최첨단 모델들은 뛰어난 추론 능력을 보여주지만, 복잡하고 현실적인 다중 단계 작업에서 어려움을 겪으며, 고객 서비스에 요구되는 높은 효율성 기준을 충족하지 못하는 경우가 많습니다. 이는 실제 기술 서비스 응용 분야에서 LLM 기반 에이전트의 향후 개발 방향에 대한 중요한 시사점을 제공합니다. CirrusBench 평가 프레임워크는 다음 주소에서 확인할 수 있습니다: https://github.com/CirrusAI

Original Abstract

The increasing agentic capabilities of Large Language Models (LLMs) have enabled their deployment in real-world applications, such as cloud services, where customer-assistant interactions exhibit high technical complexity and long-horizon dependencies, making robustness and resolution efficiency critical for customer satisfaction. However, existing benchmarks for LLM-based agents largely rely on synthetic environments that fail to capture the diversity and unpredictability of authentic customer inputs, often ignoring the resolution efficiency essential for real-world deployment. To bridge this gap, we introduce CirrusBench, a novel evaluation framework distinguished by its foundation in real-world data from authentic cloud service tickets. CirrusBench preserves the intricate multi-turn logical chains and realistic tool dependencies inherent to technical service environments. Moving beyond execution correctness, we introduce novel Customer-Centric metrics to define agent success, quantifying service quality through metrics such as the Normalized Efficiency Index and Multi-Turn Latency to explicitly measure resolution efficiency. Experiments utilizing our framework reveal that while state-of-the-art models demonstrate strong reasoning capabilities, they frequently struggle in complex, realistic multi-turn tasks and fail to meet the high-efficiency standards required for customer service, highlighting critical directions for the future development of LLM-based agents in practical technical service applications. CirrusBench evaluation framework is released at: https://github.com/CirrusAI

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!