Skip to content
Review

Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis

Aug 2026 · 0 citations · 61 references
Computer Science

TL;DR

DiagGuard is operationalized as DiagGuard, a two-stage defense-in-depth architecture in which grounding surveys available observations before localization and verification audits the diagnosis against them, showing that trajectory-level evaluation exposes limitations hidden by final-answer metrics and provides actionable guidance for improving automated RCA.

Abstract

Existing evaluations of automated root cause analysis (RCA) for microservices assess diagnostic performance mainly by endpoint correctness: whether a method localizes the responsible service. This criterion enables comparison but does not reveal the evidentiary basis of a diagnosis or the fault-propagation route connecting the source to observed symptoms, both of which an on-call site reliability engineer needs to judge whether action is warranted. We therefore treat RCA as an observable diagnostic process. Our trajectory-level framework evaluates agent executions against manually curated service-level fault-propagation paths. Applied to a public microservice RCA benchmark, it analyzes 3,500 diagnostic trajectories, characterizing where agents investigate and how they use retrieved telemetry. We find a disconnect between answer correctness and diagnostic quality: an agent may localize the fault source yet fail to reconstruct its propagation. Successful investigations stay on the fault-impact surface, act on retrieved evidence, and broaden their query repertoire as the search deepens. Failures arise when decisive evidence is omitted, retrieved evidence is misinterpreted, or unsupported inference substitutes for missing evidence. We operationalize this taxonomy as DiagGuard, a two-stage defense-in-depth architecture in which grounding surveys available observations before localization and verification audits the diagnosis against them. In an independent setting with a different model, benchmark, and service topology, DiagGuard raises Acc@1 from 43.5% to 52.5%. These results show that trajectory-level evaluation exposes limitations hidden by final-answer metrics and provides actionable guidance for improving automated RCA.

View source

Similar papers

#small language model Preprint Sep 2026

EviRCA: Decoupling Evidence Extraction from Reasoning for Microservice Root-Cause Analysis

EviRCA is presented, a framework for LLM-based RCA that decouples deterministic evidence extraction from LLM reasoning and substantially outperforms prior OpenRCA baselines that achieve up to 15.2%, while reducing token consumption by 15-26x and execution time by 3-20x.

Yu-Hao Wang, Zhen Qin, Xing-Liang Wang et al. · 0 citations
Preprint Aug 2026

ORCA: Observability-Grounded Program Repair for Microservice Incidents

Results show that operational telemetry can be transformed from diagnostic evidence into actionable repair context: paired telemetry supports repair-oriented localization, while repair graph agents convert localized code and configuration evidence into constrained patch-generation context for the LLM.

Yuan-Chen Gao, Yi-Fang Tian, Yi-Ran Li et al. · 1 citation
#machine learning Preprint Sep 2026

ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

Existing agent benchmarks mainly evaluate final task success or tool-call correctness, providing limited insight into whether agents can reliably diagnose and recover from intermediate execution failures. This limitation becomes particularly critical in multi-turn parallel tool-use scenarios, where errors may propagate...

Bo-Wen Guan, Zhen-Tao Yin, Yanming Shen · 0 citations
Preprint Aug 2026

L ONG RCA B ENCH : D IAGNOSING R ESPONSIBLE R OLES AND R OOT C AUSES IN L ONG -H ORIZON A GENT F AILURES

This work introduces LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors, and presents Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions.

Yun-Fei Zhang, Bo-Yu Feng, Changhua Pei et al. · 1 citation
Open access Oct 2026

Branch-Level Fault Localization in ADS Planning via Temporal Coverage Analysis

Planning failures in Automated Driving Systems (ADS) are increasingly detected through simulation-based testing, yet localizing their root causes within planning code remains a major challenge. Planning modules execute complex rule-based decision logic over hundreds of frames in a closed-loop interaction with the envir...

Sang-Min Woo, Dohyun Kim, Donghwan Shin et al. · 0 citations
Preprint Aug 2026

LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories

In long agent executions, an early error can persist through later actions and checks, while evidence needed to trace its origin is dispersed across the history. Short histories offer limited tests of recovering error origins across substantial subsequent execution. We introduce LongRCA Bench: 1,140 complete failed tra...

Yun-Fei Zhang, Bo-Yu Feng, Changhua Pei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.