Skip to content
Open access

Hierarchical Temporal Runtime Assurance for Controlled Agentic AI Systems: Safety Shielding, Auditable Action Repair, and Bounded Recovery

Sep 2026 · Machine Learning and Knowledge Extraction · 0 citations · 18 references

TL;DR

This study introduces MARIS-TRA, a hierarchical temporal runtime assurance extension of Controlled Agentic AI Systems that provides quantitative robustness for speed, arena, pairwise separation, restricted-zone, and bounded-recovery requirements.

Abstract

Agentic artificial intelligence requires safeguards that remain effective across trajectories rather than only at individual decisions. This study introduces MARIS-TRA, a hierarchical temporal runtime assurance extension of Controlled Agentic AI Systems. A compact, auditable fragment of Signal Temporal Logic (STL) provides quantitative robustness for speed, arena, pairwise separation, restricted-zone, and bounded-recovery requirements. The final architecture uses invariant hard-safety formulas to define feasibility, while recovery/liveness is monitored separately and may trigger escalation. Across a 5750-episode main campaign, MARIS-TRA achieved 100.00% realized hard-safety window satisfaction, 100.00% hard-safety episode satisfaction (450/450; Wilson 95% CI 99.15–100.00%), zero collision episodes, and 2.55 ms mean latency in the core comparison. Under the primary strict numerical semantics, an independent CBF-QP baseline achieved 88.89% hard-safety episode satisfaction; all 50 strict failures were very small arena-boundary overshoots in the boundary-stress scenario, with no collision or separation failures, and the post hoc tolerance sensitivity reached 100% at epsilon = 10−4 normalized simulator units. An additional 8640-episode targeted validation examined recovery hysteresis, model mismatch, AHO control flow, and safety–recovery conflicts. Under confirmatory high-density testing, safety-only shielding preserved hard safety in 810/810 episodes, whereas joint-hard enforcement produced 28/810 separation-safety failures (3.46%; Wilson 95% CI 2.40–4.95%) without collisions. Actuation-noise and delay experiments further show that the formal result is a conditional predicted-trace certification rather than a disturbance-robust guarantee on future receding-horizon execution. The empirical claims are, therefore, limited to the evaluated continuous-action multi-agent setting, while the architecture remains policy-separable and auditable.

Read PDF

Similar papers

Open access Sep 2026

Bounded Autonomy and Verifiable Safety for Agentic AI Enabled Automation

Agentic AI-enabled automation cannot be safely deployed in high-stakes environments on probabilistic reasoning alone. A recurring risk is epistemic drift: as reasoning deepens, system behavior may move away from subject-matter-expert constraints for safe operation. This paper presents BRaVeS, a bounded reasoning and sa...

S. Ramaswamy, Deveeshree Nayak · 0 citations
#artificial intelligence Preprint Sep 2026

Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems

This work formalizes the admission calculus and the assumptions connecting it to mediated execution, and establishes tested implementation behaviors and local costs, not production failure rates or comparisons of language-model capability.

Jun-Fei He, De-Ying Yu · 2 citations
Preprint Sep 2026

Regret Dominates Surprise: Design-Time Requirements Engineering for Agentic-AI Safety

A novel Regret-Dominance Mechanism (MS-RGR) is introduced to operationalize safe autonomy in agentic-AI systems, and is positioned as initial feasibility evidence for design-time safety constraints in agentic-AI requirements engineering.

Nuwayyir Almohammadi, Rami Bahsoon, Tao-An Chen · 0 citations
#machine learning Preprint Sep 2026

PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framewo...

Ding Jia, Wei Liu, Xiang-Long Du et al. · 0 citations
Preprint Aug 2026

AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit

Aviation knowledge question answering cannot directly assess the operational effectiveness and safety compliance of large language models throughout aircraft emergency procedures. We introduce AeroCopilotBench and its executable cockpit environment, ACOE, which define state-transition rules, task goals, and trajectory-...

Yu-Chen Yuan, Zheng-Huang Wu, Yuan-Gan Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.