This study introduces MARIS-TRA, a hierarchical temporal runtime assurance extension of Controlled Agentic AI Systems that provides quantitative robustness for speed, arena, pairwise separation, restricted-zone, and bounded-recovery requirements.
Abstract
Agentic artificial intelligence requires safeguards that remain effective across trajectories rather than only at individual decisions. This study introduces MARIS-TRA, a hierarchical temporal runtime assurance extension of Controlled Agentic AI Systems. A compact, auditable fragment of Signal Temporal Logic (STL) provides quantitative robustness for speed, arena, pairwise separation, restricted-zone, and bounded-recovery requirements. The final architecture uses invariant hard-safety formulas to define feasibility, while recovery/liveness is monitored separately and may trigger escalation. Across a 5750-episode main campaign, MARIS-TRA achieved 100.00% realized hard-safety window satisfaction, 100.00% hard-safety episode satisfaction (450/450; Wilson 95% CI 99.15–100.00%), zero collision episodes, and 2.55 ms mean latency in the core comparison. Under the primary strict numerical semantics, an independent CBF-QP baseline achieved 88.89% hard-safety episode satisfaction; all 50 strict failures were very small arena-boundary overshoots in the boundary-stress scenario, with no collision or separation failures, and the post hoc tolerance sensitivity reached 100% at epsilon = 10−4 normalized simulator units. An additional 8640-episode targeted validation examined recovery hysteresis, model mismatch, AHO control flow, and safety–recovery conflicts. Under confirmatory high-density testing, safety-only shielding preserved hard safety in 810/810 episodes, whereas joint-hard enforcement produced 28/810 separation-safety failures (3.46%; Wilson 95% CI 2.40–4.95%) without collisions. Actuation-noise and delay experiments further show that the formal result is a conditional predicted-trace certification rather than a disturbance-robust guarantee on future receding-horizon execution. The empirical claims are, therefore, limited to the evaluated continuous-action multi-agent setting, while the architecture remains policy-separable and auditable.
Agentic AI-enabled automation cannot be safely deployed in high-stakes environments on probabilistic reasoning alone. A recurring risk is epistemic drift: as reasoning deepens, system behavior may move away from subject-matter-expert constraints for safe operation. This paper presents BRaVeS, a bounded reasoning and sa...
S. Ramaswamy, Deveeshree Nayak· Journal of Intelligent and R...· 0 citations
This work formalizes the admission calculus and the assumptions connecting it to mediated execution, and establishes tested implementation behaviors and local costs, not production failure rates or comparisons of language-model capability.
A novel Regret-Dominance Mechanism (MS-RGR) is introduced to operationalize safe autonomy in agentic-AI systems, and is positioned as initial feasibility evidence for design-time safety constraints in agentic-AI requirements engineering.
This work introduces a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mechanically checked.
Albert Wu, N. Roberts, Tzu-Heng Huang et al.· 0 citations
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framewo...
Ding Jia, Wei Liu, Xiang-Long Du et al.· 0 citations
Aviation knowledge question answering cannot directly assess the operational effectiveness and safety compliance of large language models throughout aircraft emergency procedures. We introduce AeroCopilotBench and its executable cockpit environment, ACOE, which define state-transition rules, task goals, and trajectory-...
Yu-Chen Yuan, Zheng-Huang Wu, Yuan-Gan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.