Skip to content

Skynet: Workflow-Level Anomaly Detection for Agentic AI via Semantic and Structural Modeling

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

It is argued that anomaly detection for agentic AI must reason at the workflow level, where global execution structure exposes signals that local checks cannot see, and presents Skynet, a principled workflow-level anomaly detection framework that turns observed multi-agent execution into directed workflow graphs and scores them against learned benign behavior.

Abstract

Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these systems often originate from a single step, such as an injected prompt or a flawed plan, and are then amplified through downstream dependencies as the corrupted step propagates across many subsequent agents and tool calls. Existing defenses either target a specific class of attacks or failures, or inspect individual prompts and steps in isolation. Both leave the global dependency structure of a workflow unexamined, and miss the inconsistencies that only emerge when the execution is viewed as a whole. We argue that anomaly detection for agentic AI must reason at the workflow level, where global execution structure exposes signals that local checks cannot see. We present Skynet, a principled workflow-level anomaly detection framework that turns observed multi-agent execution into directed workflow graphs and scores them against learned benign behavior. Skynet jointly models the semantic execution context and the structural organization of inter-agent delegation, tool invocation, and data-flow dependencies, and trains only on benign workflows. Because training never sees attacks or failures, this design naturally extends to zero-day detection: any execution that violates benign workflow regularities surfaces as off-manifold geometry under a single decision rule. We evaluate Skynet on three public agentic safety and failure benchmarks. It sustains high recall together with a sub-1% false positive rate, with per-workflow and per-step latencies low enough for online monitoring of agentic AI runtimes.

View source

Similar papers

Conference Aug 2026

Reconceptualizing Observability for Agentic AI Systems: A Trace-Centric Architecture for Interpreting Non-Deterministic Workflow Behavior

The more typical feature of agentic AI systems is dynamic, multistep workflows where autonomous components plan, reason, and communicate with external tools and data sources in a series of iterations. Such flexibility increases capability but also brings nondeterminism which is inherent and where the same inputs can re...

Ankur Gupta, Karan Gupta, Divyakumar Deepak Savla et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf is a profiler that aggregates agent trajectories into pprof-compatible profiles, enabling flame graph visualization and analysis and introduces recursive operation segmentation, which recursively splits trajectories at task boundaries.

Yu-Sheng Zheng, Chaokun Chang, Yuan-Man Mao et al. · 0 citations

DeepEye: A Workflow-Centric Agentic Data System for Steerable Data Analytics

This work presents DeepEye, a workflow-centric agentic data system that turns user intents into transparent and steerable analytical workflows and develops DataMagic as the system’s Video Generator, a declarative multi-agent method that improves data-video quality.

Unknown authors · 0 citations

DeepEye: A Workflow-Centric Agentic Data System for Steerable Data Analytics

This work presents DeepEye, a workflow-centric agentic data system that turns user intents into transparent and steerable analytical workflows and develops DataMagic as the system’s Video Generator, a declarative multi-agent method that improves data-video quality.

Unknown authors · 0 citations
#artificial intelligence Preprint Oct 2026

DeFA: Dependency-Guided Failure Attribution for LLM Agents

Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines protocol relations and s...

Bo Deng, Xin-Lei Zheng, Yi-Xun Wei et al. · 0 citations
#artificial intelligence Review Sep 2026

Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

Agentic workflows now make consequential decisions in regulated settings, and the governance placed around them is almost entirely step-scoped: input-output classifiers, per turn rails, and span-level evaluators. The policies organizations actually hold, such as referral thresholds, authority limits, and review require...

Ashwini Kurady, S. Grandhi, R. Gupta et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.