Multi-Tenant GPU Clouds: Networking Patterns That Let Many Teams Safely Share One AI Factory

Authors

  • Indra Kumar Mondal Independent Researcher, USA

DOI:

https://doi.org/10.70917/ijcisim-2026-5514

Keywords:

multi-tenant GPU clusters, network isolation, GPU virtualization, noisy-neighbor interference, network-aware scheduling, cloud security

Abstract

Purpose: Large graphics processing unit (GPU) clusters, once built as single-purpose research infrastructure, are now routinely operated as shared, multi-tenant "AI factories" serving many independent teams on the same physical hardware. This introduces three risk classes that cloud multi-tenancy literature and GPU infrastructure literature have largely treated as separate problems: performance interference between co-located tenants, security exposure from hardware co-residency, and unfair resource allocation. This paper examines whether isolation mechanisms developed independently at the network layer, the GPU layer, and the scheduler layer are individually sufficient, or whether they must be coordinated as a single cross-layer problem.

Design/methodology/approach: This is a literature-based comparative analysis. A three-layer taxonomy organizes the review: network-layer isolation (virtual network overlays, software-defined slicing, transport-level quality-of-service), GPU-layer partitioning (hardware virtualization and process-level sharing), and scheduler-layer fairness (network-aware, congestion-aware job placement). Evidence is drawn from foundational multi-tenant networking research, empirical production-cluster characterizations, recent GPU-specific security research, and the most recent published giga-scale AI fabric architecture descriptions.

Findings: Network-layer isolation, while mature for general cloud multi-tenancy, cannot address risks arising inside the GPU and its interconnect. GPU-layer partitioning substantially reduces performance interference but introduces its own security surface, including GPU-specific covert and side-channel attacks demonstrated on shared systems as recently as 2025. Scheduler-layer fairness mechanisms accounting for network topology and congestion state measurably reduce interference in production-scale clusters, but depend on signals from both other layers. No single layer's isolation is sufficient alone; safe multi-tenancy requires coordinated isolation across all three.

Originality/value: This paper's contribution is the three-layer taxonomy itself, connecting network-layer, GPU-layer, and scheduler-layer isolation research that has developed largely independently across different research communities. It identifies GPU-specific security isolation as the least mature layer and the area most likely to determine whether current multi-tenant GPU cloud architectures remain safe as adversarial research matures.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-04

How to Cite

Indra Kumar Mondal. (2026). Multi-Tenant GPU Clouds: Networking Patterns That Let Many Teams Safely Share One AI Factory. International Journal of Computer Information Systems and Industrial Management Applications, 18(22s), 1234–1240. https://doi.org/10.70917/ijcisim-2026-5514

Issue

Section

Original Articles