AI-Enabled Cloud Resource Management for Improving Quality of Service and Reducing Computational Cost
DOI:
https://doi.org/10.70917/ijcisim-2026-5839Keywords:
Cloud computing, resource management, auto-scaling, workload forecasting, gradient boosting, quality of service, cost optimisationAbstract
Cloud data centres must honour strict Quality of Service (QoS) commitments while the cost of leased and powered computing capacity keeps rising. Rule-based elasticity mechanisms react only after demand has already changed, forcing operators to choose between wasteful over-provisioning and SLA-breaching under-provisioning. This paper presents an AI-enabled resource management framework organised as a monitor–forecast–decide–act loop. A gradient-boosted regression model predicts the request rate one provisioning delay ahead, and a cost-aware controller converts the forecast into a virtual-machine (VM) target through a validation-derived safety margin, an online residual correction and hysteresis on scale-in. The framework is evaluated in a discrete-time simulator on ten independent synthetic workloads containing diurnal, weekly and flash-crowd components, and is compared with static peak provisioning, a reactive threshold auto-scaler and a linear-regression predictive scaler. The proposed controller lowered daily cost by 15.9% relative to the reactive policy and by 54.8% relative to static provisioning, matched the reactive SLA-violation rate (1.05% versus 1.04%) and reduced 95th-percentile latency by 16.8%. The study also documents where the approach falls short, notably during bursts that exceed the range seen in training.