CausalDX: Diagnosing Long-Tail and Cascading Cloud Incidents With LLM-Guided Causal Reasoning
Abstract
Modern cloud platforms face escalating diagnostic challenges due to scale-driven emergent behaviors and intricate system interdependencies. However, contemporary diagnostic systems can only handle well-understood failure patterns, leaving two critical blind spots that impede system reliability: (1) <italic>Long-tail diagnostic paralysis</italic>, where rare anomalies evade pattern-based detection, consuming over 60% of diagnostic time despite their small proportion, and (2) <italic>Cascading Anomaly Overload</italic>, where exponentially large anomaly combinations induce spurious correlations, making it challenging to distinguish between surface-level symptoms and root causes. To bridge these gaps, we introduce <sc>CausalDX</sc>, a novel diagnostic framework that leverages LLMs as causal inference engines to extend the generalization capabilities of diagnostic systems. Our framework consists of three key components: (1) an anomaly-granular causal graph representation that structures incident diagnosis as a systematic graph traversal process, enabling clear and explainable reasoning; (2) the Adaptive Graph-based Root Cause Search (AGRCS) algorithm that combines expert rules with LLM agents to decompose complex root-cause analysis into manageable steps; and (3) an observation-based self-verification mechanism that uses variational inference to validate root-cause confidence against actual system observations. Extensive evaluation on real-world cloud incidents shows that <sc>CausalDX</sc> achieves significant improvements: a 73.2% automated diagnostic success rate, with 35.2% recall for long-tail anomalies (up to 3× over LLM-based SOTA) and 68.4% precision in cascading scenarios (<inline-formula><tex-math notation="LaTeX">$3.2\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>3</mml:mn><mml:mo>.</mml:mo><mml:mn>2</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="cui-ieq1-3722256.gif"/></alternatives></inline-formula> over current systems). The verification mechanism reduces hallucination-induced false positives by 17.9%. These results show that our design effectively harnesses LLMs for adaptive, explainable, and reliable cloud incident management.