Skip to content

CausalDX: Diagnosing Long-Tail and Cascading Cloud Incidents With LLM-Guided Causal Reasoning

Oct 2026 · IEEE Transactions on Knowledge and Data Engineering · Vol 38, pp. 6424-6438 · 0 citations · 75 references

Abstract

Modern cloud platforms face escalating diagnostic challenges due to scale-driven emergent behaviors and intricate system interdependencies. However, contemporary diagnostic systems can only handle well-understood failure patterns, leaving two critical blind spots that impede system reliability: (1) <italic>Long-tail diagnostic paralysis</italic>, where rare anomalies evade pattern-based detection, consuming over 60% of diagnostic time despite their small proportion, and (2) <italic>Cascading Anomaly Overload</italic>, where exponentially large anomaly combinations induce spurious correlations, making it challenging to distinguish between surface-level symptoms and root causes. To bridge these gaps, we introduce <sc>CausalDX</sc>, a novel diagnostic framework that leverages LLMs as causal inference engines to extend the generalization capabilities of diagnostic systems. Our framework consists of three key components: (1) an anomaly-granular causal graph representation that structures incident diagnosis as a systematic graph traversal process, enabling clear and explainable reasoning; (2) the Adaptive Graph-based Root Cause Search (AGRCS) algorithm that combines expert rules with LLM agents to decompose complex root-cause analysis into manageable steps; and (3) an observation-based self-verification mechanism that uses variational inference to validate root-cause confidence against actual system observations. Extensive evaluation on real-world cloud incidents shows that <sc>CausalDX</sc> achieves significant improvements: a 73.2% automated diagnostic success rate, with 35.2% recall for long-tail anomalies (up to 3× over LLM-based SOTA) and 68.4% precision in cascading scenarios (<inline-formula><tex-math notation="LaTeX">$3.2\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>3</mml:mn><mml:mo>.</mml:mo><mml:mn>2</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="cui-ieq1-3722256.gif"/></alternatives></inline-formula> over current systems). The verification mechanism reduces hallucination-induced false positives by 17.9%. These results show that our design effectively harnesses LLMs for adaptive, explainable, and reliable cloud incident management.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.