Skip to content
Open access

EviGuard: Machine-Verifiable Evidence Grounding for LLM-Based Industrial Incident Reasoning

Sep 2026 · Applied Sciences · Vol 16, pp. 8925 · 0 citations · 25 references

TL;DR

EviGuard is presented, a system that decides when an LLM’s understanding is trustworthy enough to act on and has an ensemble of deterministic verifiers label every claim supported, contradicted, or unknown against the graph—honoring interval time, event-time policy and credential versions, network reachability, and physical control dependencies.

Abstract

Large language models (LLMs) can turn a flood of cross-layer industrial logs into a fluent incident narrative, but a narrative that cites only real, resolvable events can still be wrong in every relation that matters: the login came from a different workstation, the write command occurred after the physical change it supposedly caused, the action fell inside a planned maintenance window, and the controller does not even actuate the affected process. A cited event is not necessarily supporting evidence. When such a narrative drives automated response, the error propagates into isolating the wrong controller or revoking a legitimate operator. We present EviGuard, a system that decides when an LLM’s understanding is trustworthy enough to act on. EviGuard stores auditable cross-layer evidence in a provenance graph, lets the LLM propose only hypotheses, compiles each hypothesis into atomic machine-checkable claims in an Incident Claim Language, and has an ensemble of deterministic verifiers label every claim supported, contradicted, or unknown against the graph—honoring interval time, event-time policy and credential versions, network reachability, and physical control dependencies. A response gate forbids any high-impact action whose critical preconditions are not all supported. On EviCPS-Bench (42 hardware-in-the-loop attack chains, 9600 claim-level labels, κ=0.87), EviGuard cuts the unsupported-claim rate from 12.6% to 1.7%, raises relation-edge F1 from 0.64 to 0.89, holds prompt-injection success to 0.4%, and executes zero unverified high-impact actions across 3200 response decisions, at a median end-to-end latency of 0.44 s.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Janus: Evidence-Before-Effect Sagas and Offline-Verifiable Provenance for Agentic LLMs

Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. Janus puts the record on the effect path. A step's proposal, the verdict on it and any answer from a validator or a person are durable in a signed, hash-chaine...

Mustafa Arslan · 0 citations
Preprint Aug 2026

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, answer-consistent, and auditable.

Srimonti Dutta, Akshata Kishore Moharir · 0 citations
Review Sep 2026

Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents

Autonomous coding agents read untrusted files, run shell commands and spawn sub-agents with little supervision, yet their record is usually an editable log. We present Tracekit, an open-source, dependency-free system that captures three channels for every agent session: what the human asked (intent), what the model sai...

Bravish Ghosh · 0 citations
#machine learning Preprint Oct 2026

Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning

Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to o...

Boyue Caroline Hu, K. Ahir, Ronghao Ni et al. · 0 citations
Review Sep 2026

Guarded Commits: Transactional Human Approvals for LLM Workflows

LLM workflows often require human approval before an irreversible external action. Most systems keep that approval outside the workflow, as an interface click or an audit entry. The workflow therefore lacks a commit-time check that every risky path reached an approval gate. Its logs may not preserve the reviewed eviden...

Laurent Bindschaedler, Ferdinand Kossmann, Chun-Wei Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.