Results show that operational telemetry can be transformed from diagnostic evidence into actionable repair context: paired telemetry supports repair-oriented localization, while repair graph agents convert localized code and configuration evidence into constrained patch-generation context for the LLM.
Abstract
Microservice failures are often diagnosed from operational telemetry. However, automated program repair systems usually start from issue reports, localized code context, or failing tests. This mismatch leaves a gap between telemetry-based diagnosis and patch generation. We present ORCA, an observability-grounded APR pipeline for microservice incidents. ORCA first distills the differences in paired failure and reference telemetry into a fault signature, then uses the signature to identify candidate code and deployment-configuration locations. Repair graph agents and an Exploration agent generate unified-diff patch candidates from these locations. ORCA evaluates generated patches with a Telemetry-Grounded Patch Verifier that separates patch validity, syntactic and semantic correctness, test-oracle integrity, and telemetry replay. On a 575-case benchmark, ORCA outperforms all evaluated baselines in terms of cost-effectiveness. Results show that operational telemetry can be transformed from diagnostic evidence into actionable repair context: paired telemetry supports repair-oriented localization, while repair graph agents convert localized code and configuration evidence into constrained patch-generation context for the LLM. Telemetry-grounded verification then exposes repair outcomes that issue- or test-only evaluation would miss.
Microservice-based systems are tightly interdependent. A failure in one service can propagate to a dependent service, potentially bringing down critical, user-facing functionality. We have developed and deployed RPCShield, a suite of program analyses for Go and Java that detects such cascading-failure risks. RPCShield...
Milind Chabbi, Sonal Mahajan, Ivan Beschastnikh et al.· Proceedings of the ACM SIGOP...· 1 citation
DiagGuard is operationalized as DiagGuard, a two-stage defense-in-depth architecture in which grounding surveys available observations before localization and verification audits the diagnosis against them, showing that trajectory-level evaluation exposes limitations hidden by final-answer metrics and provides actionab...
Qisheng Lu, Ao-Yang Fang, Junjielong Xu et al.· 0 citations
IntentP4 is presented, a formal-methods-aided pipeline that translates an operator's natural-language intent into a P4LTL specification and then into a replayable multi-packet test case, grounded throughout in compiler artifacts via a tool-queryable ProgramContext and gated by deterministic per-stage validators.
Ruonan Feng, Mingming Zhang, Yu Jiang et al.· Conference on Applications,...· 0 citations
Modern enterprise cloud infrastructures need to be able to deliver software quickly, remain operationally resilient and follow strict regulatory requirements. Through traditional operations, observability, security and governance are typically treated as separate processes, which leads to longer Mean Time To Repair (MT...
Vishwa Lakhnakiya· 2026 International Conferenc...· 0 citations
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We presen...
Yi-Fang Tian, Ying-Jian Bai, Yi-Feng He et al.· 0 citations
Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas, are introduced, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas.
Sugam Panthi, Rabab Abdelfattah· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.