Skip to content
Open access

OpenReliOps: A Safety-Constrained Architecture for Autonomous Reliability in Open-Weight Model Services

Prudvi Saisaran Ponduru Pavani Priya Vyshnavi Nandanavanam Sai Kesav Kumar Ponduru
Aug 2026 · International Journal of Advanced Multidisciplinary Research and Studies · Vol 6, pp. 1114-1121 · 0 citations

TL;DR

OpenReliOps is presented, a self-healing control architecture that treats an open-weight model as a replaceable reasoning component inside an evidence-grounded and policy-bounded reliability loop and is positioned as a falsifiable methodological contribution.

Abstract

Open-weight language models enable private deployment, local adaptation, and independent audit, but they also transfer responsibility for model artifacts, inference infrastructure, telemetry, security, and generated actions to the deploying organization. This paper presents OpenReliOps, a self-healing control architecture that treats an open-weight model as a replaceable reasoning component inside an evidence-grounded and policy-bounded reliability loop. The architecture integrates service-level-objective forecasting, topology-aware telemetry retrieval, causal root-cause diagnosis, a typed repair domain-specific language, independent policy verification, calibrated escalation, progressive execution, rollback, and post-action semantic and operational validation. A reliability process preference optimization objective is defined to learn from successful, rejected, failed, and rolled-back incident trajectories without rewarding unsafe outcome-only behavior. Formal analysis establishes policy confinement and blast-radius bounds under complete mediation, current state, typed actions, and sound declared policies. The evaluation protocol separates diagnostic correctness, evidence grounding, policy safety, and measured service recovery, and specifies comparisons on RCAEval, OpenRCA, OpenRCA 2.0, AIOpsLab, and controlled model-serving faults. The analytical result is that model substitution alone cannot provide autonomous reliability: the reliability boundary is created by the interfaces among evidence, uncertainty, policy, execution, and verification. The framework is therefore positioned as a falsifiable methodological contribution whose operational gains must be demonstrated through incident-level experiments, ablations, multiple seeds, and complete failure reporting.

Read PDF

Similar papers

Open access Sep 2026

Telemetry-Native Reliability Engineering for Agentic AI Systems: Topology-Aware Failure Correlation and Decision Traceability

Agentic artificial intelligence systems combine large language models, tool-calling runtimes, retrieval services, application programming interfaces, policy guardrails, memory stores, and human approval paths. These systems create a reliability problem that differs from conventional microservices because failures may a...

Shivam Dube, R. Varshney, S. Kumar · 0 citations
#artificial intelligence Preprint Sep 2026

Autonomy in Check: Governor-Mediated Adaptive Security at the Edge

Adaptive security at the network edge increasingly relies on automated planners, including rule-based controllers, learned policies, and LLM-assisted agents, that translate observations into enforcement actions. Once such a planner can influence live policy state, syntactic validity is not enough. A semantically wrong...

Ijaz Ahmad, Flavio Esposito, Erkki Harjula · 0 citations
Open access Sep 2026

A Responsibility-Allocation Framework for LLM-First and Hybrid Code-First Enterprise AI Architectures

Enterprise use of large language models creates architectural challenges involving control, traceability, state management, output conformance, and operational governance. This design-science paper introduces the Deterministic– Probabilistic Responsibility Allocation Framework, which distinguishes a model-first archite...

Swapneswar Sundar Ray · 0 citations
Preprint Aug 2026

TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions

TwinGridShield is presented, a model-independent runtime authorization layer that evaluates each proposed action in a deterministic network twin before release that verifies conformance of the implementation to its encoded authorization predicate rather than safety under model error.

M. Rafy · 0 citations
Review Open access Sep 2026

LLMOps: A Foundation Model–Driven Framework for Autonomous Cloud Reliability Engineering

This research introduces Large Language Model Operations (LLMOps), an innovative Foundation Model–Driven Autonomous Cloud Reliability Engineering Framework that incorporates Large Language Models (LLMs) throughout the entire cloud operations lifecycle, thereby providing a scalable foundation for next-generation autonom...

M. Dhanekula · 0 citations
#artificial intelligence Preprint Sep 2026

Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems

This work formalizes the admission calculus and the assumptions connecting it to mediated execution, and establishes tested implementation behaviors and local costs, not production failure rates or comparisons of language-model capability.

Jun-Fei He, De-Ying Yu · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.