Skip to content
Open access

From Reactive Pipelines to Self-Healing Data Platforms: An Agentic AI Framework for Reliable, Secure, and Cost-Efficient Azure Data Engineering

Sep 2026 · American Journal of Technology · 0 citations

TL;DR

The study supports the use of bounded, auditable agentic control for recurring operational failure modes, and suggests that policy-gated autonomy lowers mean time to resolve (MTTR) and incident recurrence.

Abstract

Aim: This study aimed to design, implement, and empirically evaluate an agentic AI framework that improves the reliability, governance, security, and cost efficiency of enterprise Azure data platforms. The framework was intended to move operations from reactive, manual incident handling to policy-constrained automated monitoring and remediation, while preserving auditability, least privilege, and human oversight for high-risk actions. Specifically, the study sought to determine whether bounded agentic control could reduce operational toil, improve pipeline success and recovery times, strengthen security-governance posture, and lower unit costs without violating service-level or compliance constraints. Methods: We propose an agentic AI framework that (1) continuously telemetries pipeline runs, data-quality checks, lineage, and security posture; (2) retrieval-augments reasoning on operational knowledge (tickets, runbooks, KQL logs, IaC diffs); (3) policy-constrained action execution (RBAC, approvals, change windows, least privilege) to remediate failures, enforce baselines, and optimize resources; and (4) post-action validation to confirm recovery and prevent regressions. The system was built on Azure OpenAI and Azure Databricks, Data Factory, and Microsoft Fabric and tested with an enterprise deployment, a historical incident replay, and an A/B test against standard on-call procedures. Results: Manual interventions decreased by 65% across workloads, pipeline success rate increased from 91% to >97%, and annualized savings approached USD $1M, with better security-governance scores and lower cost per successful run. Results suggest that policy-gated autonomy lowers mean time to resolve (MTTR) and incident recurrence. Conclusion: The study supports the use of bounded, auditable agentic control for recurring operational failure modes. Recommendation: Future work should strengthen robustness guarantees, standardize multi-objective evaluation, and assess portability beyond Azure.

Read PDF

Similar papers

Open access Sep 2026

Security Architecture for Agentic AI in Enterprise Cloud Environments: A Zero-Trust Framework for Secure Autonomous Systems

Agentic artificial intelligence expands the enterprise security boundary because autonomous agents can plan tasks, retain memory, invoke tools, call APIs, and initiate business actions. Authentication at session start is therefore insufficient when later actions may be influenced by untrusted content, poisoned memory,...

S. Suryawanshi · 0 citations
#artificial intelligence Review Sep 2026

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and operational risks. Untrusted content from web...

Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh et al. · 0 citations
Review Open access Sep 2026

Artificial Intelligence - Model Context Protocol - Review of Real-World Security Threats

Modern Artificial Intelligence (AI) infrastructures and Model Context Protocol (MCP) deployments face systemic security exposures as a result of autonomous agents interacting directly with external databases and tools. Protocol-level weaknesses, permissive access rights, and unverified third-party repositories allow at...

Upendra Kanuru, Alexa Schmitt · 0 citations
2026

Governing Agentic AI in Enterprise Operations: Architectural “Rails” for Safe, Deterministic, and Compliant Autonomous Systems

This paper argues that the introduction of agentic AI requires a substantial expansion of traditional enterprise architecture principles to address new behavioral, security, and governance risks emerging from non-deterministic AI systems interacting with heterogeneous operational platforms-ERP, HCM, CLM, asset manageme...

Elizabeth Koumpan, Vimal Dimpi · 0 citations
Preprint Sep 2026

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a com...

Han-Zhang Ma, Alicia Y. Hariri, Tian-Xiang Shen et al. · 0 citations
Review Open access Aug 2026

A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security

Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agen...

Sanket Badhe, D. Shah, Priyanka Tiwari et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.