Toward Self-Healing Cloud-Native Infrastructure for Production Large Language Model Workloads

Authors

  • Rajeev Chevuri Campbellsville University, USA

DOI:

https://doi.org/10.70917/ijcisim-2026-5799

Keywords:

Self-healing infrastructure, cloud-native systems, Kubernetes orchestration, large language model operations, automated remediation, observability

Abstract

Production large language model (LLM) infrastructure is a cloud-native system that combines container orchestration, accelerated compute, model-serving runtimes, storage, networking, and deployment pipelines into a single operational surface, and a fault in any one layer can degrade the availability of the whole inference service. Reactive operations, in which monitoring raises an alert and an engineer manually diagnoses and resolves the condition, remain necessary for genuinely novel failures but scale poorly against the recurring, well-understood conditions that dominate day-to-day LLM operations. This paper proposes a structured alternative: a seven-domain failure taxonomy spanning the container, scheduling, node, accelerator, networking, storage, provisioning, and deployment-pipeline layers, paired with a six-layer reference architecture, instrumentation, correlation, classification, policy, execution, and validation and escalation, that governs when automated remediation is permitted and when a condition must reach a human operator. The architecture treats observability as a decision layer rather than a passive monitoring function, combining signals across layers to establish operational state and event correlation before any corrective action executes. Remediation itself is bounded: actions carry retry limits, validation checks, and reversibility requirements, and the architecture explicitly addresses how automation can itself introduce instability through repeated retry loops, conflicting controllers, or premature intervention during transient conditions. The paper argues that self-healing must be designed as an architectural capability from the outset rather than retrofitted after reliability problems emerge, and that automation quality, measured by whether actions are bounded, validated, and reversible, matters more than automation coverage. The resulting model gives infrastructure teams a reusable basis for scaling LLM operations without either relying entirely on manual intervention or ceding recovery decisions to unrestricted autonomous systems.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-03

How to Cite

Rajeev Chevuri. (2026). Toward Self-Healing Cloud-Native Infrastructure for Production Large Language Model Workloads. International Journal of Computer Information Systems and Industrial Management Applications, 18(22s), 1919–1930. https://doi.org/10.70917/ijcisim-2026-5799

Issue

Section

Original Articles