CTBench: 실제 통신망 운영 환경에서 AI 에이전트의 문제 해결 능력 평가
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
에이전트는 네트워크 운영 및 유지 관리 자동화를 위해 점점 더 많이 활용되고 있으며, 엔지니어는 엄격한 제약 조건 하에서 네트워크 오류를 진단하고, 서비스 향상을 위한 구성을 최적화하며, 운영 비용을 절감해야 합니다. 그러나 기존의 평가 방법은 실제 네트워크 특성을 정확하게 반영하지 못하거나, 다양한 벤더, 장치, 프로토콜 및 인터페이스를 가진 부분적으로 관찰 가능한 통신 환경에서 에이전트를 평가하는 데 한계가 있습니다. 본 논문에서는 에이전트가 숙련된 통신 문제 해결 엔지니어처럼 작동하는지 평가하기 위한 공개 벤치마크인 CTBench를 소개합니다. CTBench는 근본 원인 분석 및 경로 복구에 중점을 두고 있으며, 각 작업은 전문가에 의해 구성되고 '황금' 증거 단계를 포함한 풍부한 메타데이터로 주석이 달려 있습니다. CTBench는 최종 답변과 진단 증거 모두를 평가하는 전문가 기반의 지표를 사용합니다. 대표적인 하드웨어-모델 조합을 사용한 실험 결과, 최첨단 에이전트는 경로 복구 작업에서 엔드포인트를 식별하는 데 매우 뛰어난 성능을 보이지만, 전반적으로 근본 원인 분석에서는 성능이 저조했습니다. 특히, 에이전트는 인터페이스 상태, 링크 계층, 서비스 관리 및 기타 운영 관련 오류를 해결하는 데 어려움을 겪습니다. 더욱 중요한 점은, 에이전트가 타당하거나 정확한 최종 답변을 제시하더라도 실제 운영 환경에서 요구되는 증거 기반 진단을 제공하지 못하는 경우가 많습니다. 또한, 우리의 연구 결과는 경로 복구가 일반적으로 더 많은 리소스를 필요로 하지만, 더 큰 리소스 사용량이 반드시 더 나은 진단으로 이어지는 것은 아니라는 것을 보여줍니다.
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.