CloudCons: 클라우드 리소스 통합을 위한 종합적인 엔드투엔드 벤치마크
CloudCons: A Comprehensive End-to-End Benchmark for Cloud Resource Consolidation
서비스 안정성을 보장하기 위해 과도하게 자원을 할당하는 방식 때문에, 클라우드 데이터 센터의 자원 활용률은 여전히 낮은 수준입니다. 이를 해결하기 위해, 미래 수요를 예측하여 최적화하는 '예측 후 최적화' 패러다임이 등장했습니다. 최근에는 시계열 기반 모델들이 제로샷 일반화를 통해 이 패러다임을 향상시킬 수 있는 가능성을 보여주지만, 기존 벤치마크는 예측 오류 지표에만 초점을 맞추고 있습니다. 이러한 고급 모델들의 실제 의사 결정 유용성은 검증되지 않았으며, 따라서 다운스트림 작업에서의 실질적인 가치는 불확실합니다. 이 문제를 해결하기 위해, 우리는 클라우드 리소스 통합이라는 특정 맥락에서 예측 모델을 평가하도록 설계된 종합적인 엔드투엔드 벤치마크인 CloudCons를 제안합니다. Huawei Cloud, Microsoft Azure 및 Google Borg의 다양한 워크로드를 포함하는 고품질 데이터 세트를 구축하여 동기화된 일주기 리듬부터 확률적이고 불규칙한 패턴, 그리고 고주파 노이즈에 이르기까지 다양한 서비스 특성을 포착했습니다. 통계 모델, 딥러닝 모델 및 기반 모델을 광범위하게 평가했습니다. 우리의 실험 결과는 중요한 사실을 보여줍니다: 기반 모델은 뛰어난 제로샷 예측 정확도를 보이지만, 이러한 장점이 반드시 더 나은 의사 결정 유용성으로 이어지지는 않습니다. 실질적으로 중요한 점은 예측 분위수의 선택이 매우 중요한 요소라는 것입니다. 우리는 자원 효율성과 서비스 안정성 사이의 균형을 맞추기 위해 이러한 선택을 조정하는 방법에 대한 실행 가능한 지침을 제공하며, 이는 실제 배포 결정에 필수적인 통찰력을 제공합니다.
Driven by conservative over-provisioning to guarantee service reliability, resource utilization in cloud data centers remains at low levels. To mitigate this, the forecast-then-optimize paradigm has emerged to optimize consolidation by anticipating future demands. While emerging time series foundation models promise to enhance this paradigm through zero-shot generalization, existing benchmarks focus solely on prediction error metrics. The actual decision utility of these advanced models remains unverified, rendering their practical value for downstream tasks uncertain. To bridge this gap, we propose CloudCons, a comprehensive end-to-end benchmark designed to evaluate forecasting models within the specific context of cloud resource consolidation. We build high-quality datasets that cover diverse workloads from Huawei Cloud, Microsoft Azure, and Google Borg, capturing distinct service characteristics ranging from synchronized diurnal rhythms to stochastic, pulse-like bursts and high-frequency noise. We conduct an extensive evaluation of statistical, deep learning, and foundation models. Our experiments reveal a pivotal finding: while foundation models demonstrate superior zero-shot forecasting accuracy, this advantage does not inherently translate into better decision utility. Of practical significance, we systematically analyze how the selection of predictive quantiles acts as a critical lever. We provide actionable guidelines for calibrating these selections to balance the trade-off between resource efficiency and service reliability, offering vital insights for real-world deployment decisions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.