Scalable General-Purpose Cloud Compute: Architectural Innovations for Sustainable AI Infrastructure
DOI:
https://doi.org/10.70917/ijcisim-2026-5328Keywords:
Cloud Computing, Sustainable Infrastructure, Virtualization, Cloud Architecture, Resource Optimization, Hardware Lifecycle, Compute Express LinkAbstract
The rapid growth of artificial intelligence (AI), machine learning, and large-scale digital services is placing unprecedented demand on cloud infrastructure, making scalable and sustainable compute increasingly important. While advances in processors and accelerators continue to improve computational capability, architectural coordination can unlock efficiency gains that hardware-generation improvements alone may not fully capture, particularly across increasingly heterogeneous compute environments (Armbrust et al., 2010). This paper will examine virtualization, workload scheduling, the provisioning of heterogeneous resources and infrastructure lifecycle management's effects on infrastructure efficiency at hyperscale. The paper extends the NIST cloud service definition (Mell & Grance, 2011) and most recent research on energy efficient resource management (Khan et al., 2022, Ilager et al., 2021) to propose a Sustainable Compute Lifecycle Framework comprised of five inter-connected phases: Platform Planning, Platform Deployment and Optimization, Platform Operations, Hardware Modernization, and Hardware Retirement and Resource Reclamation. The paper also explores new trends such as Compute Express Link (CXL) for memory pooling and disaggregation (Das Sharma et al., 2024;Chen et al., 2024) and the use of AI for infrastructure scheduling (Sanjalawe et al.,2025). The key message is that there is a new central design coordination of the cloud architecture that enables sustainable growth of the cloud, not merely increasing incremental capacity on hardware.