Skip to content

Author

Ajay Devineni

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Chaos Engineering and Observability Integration for Resilient Cloud Infrastructure

Cloud computing acts as the core foundation supporting modern digital services, and hosts large-scale applications across five major domains: finance, healthcare, e-commerce, education, and industrial automation. However, its inherent complex, distributed, and dynamic native characteristics make it susceptible to four types of failures: sudden latency spikes, service outages, resource exhaustion, and security vulnerabilities. Traditional monitoring systems, which can only respond to failures after they occur, cannot guarantee the resilience of cloud environments. To address this issue, this paper proposes an integrated framework that combines chaos engineering and observability: chaos engineering injects controlled failures into production-like environments to evaluate a system’s load-bearing capacity, while observability obtains in-depth insights into a system through metrics, logs, distributed tracing, and event analysis. This framework unifies the capabilities of the two types of platforms to realize three core functions: proactive failure detection, automated recovery, and continuous resilience verification. We conducted validation experiments based on Kubernetes-powered containerized microservices, paired with Prometheus, Grafana, Jaeger, and LitmusChaos. Experimental results show that the framework achieves notable improvements across four dimensions: failure detection time, system recovery rate, service availability, and operational reliability. It can help all types of organizations identify hidden vulnerabilities, cut downtime, and strengthen service continuity.

Ganesh Gurudu, Ajay Devineni · 0 citations