Skip to content
Open access

A Cloud-Native MLOps Framework for Drift Detection, Automated Retraining, and Reliable Real-Time Inference in Predictive AI Systems

Sep 2026 · International journal of data science and machine learning · 0 citations

Abstract

Deployed predictive machine learning models inevitably degrade when real-world production data deviates from training distributions. Static deployment strategies cannot handle this environmental shift, causing drops in accuracy and forcing engineering teams to rely on slow, manual retraining audits. This paper presents an end-to-end, autonomous MLOps framework that bridges this gap by coupling continuous streaming telemetry with containerized retraining loops on Kubernetes. The architecture relies on an independent microservices design, utilizing Prometheus to track live data features and executing Kolmogorov-Smirnov and Population Stability Index (PSI) algorithms to catch statistical drift. We validated the system under a simulated fraud-detection workload streaming 50,000 requests per minute. Under testing, the framework flagged a 25% covariate shift within 2.5 minutes of its introduction. The automated pipeline completed retraining and package validation on an isolated compute plane, restoring predictive accuracy (F1-score > 0.93) within a total recovery window of 10.6 minutes. By implementing progressive canary rollouts for the live model swap, the system maintained a strict P99 real-time inference latency under 50ms with zero request drops or system downtime. These results demonstrate that tight architectural integration between telemetry and orchestration removes human-in-the-loop operational overhead and keeps live predictive systems resilient against distribution decay.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.