A fully on-premise clinical agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy is developed and evaluated, supporting a practical framework for institutionally governed clinical agents.
Abstract
Autonomous clinical artificial intelligence (AI) agents powered by large language models (LLMs), meaning systems that can complete a diagnostic workflow without continuous human input, are increasingly capable of supporting complex reasoning and decision-making. Clinical translation, however, remains limited by two unmet requirements: institutionally governed deployment and reliable decision-time uncertainty estimation. Here we developed and evaluated a fully on-premise clinical agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy. Across two Medical Information Mart for Intensive Care IV (MIMIC-IV)-derived benchmarks, the agent achieved 90.04% accuracy on a seven-disease task and 83.8% accuracy on a four-disease task, approaching a cloud baseline on the primary benchmark. To assess decision-time reliability, we quantified internal-likelihood, language-based and behavioral-stability measures for diagnosis and reasoning. Diagnostic behavioral consistency provided the strongest discrimination of correctness (area under the curve (AUC) = 0.860) and remained informative under stress testing (AUC = 0.875). At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy. These findings support a practical framework for institutionally governed clinical agents in which decision-time reliability signals identify a lower-risk subset for autonomous handling and defer the remainder for review.
Clinical artificial intelligence (AI) has advanced rapidly, with frontier large language models now matching or exceeding physician performance on simulated diagnostic reasoning and clinical decision-support tasks. Yet adoption has outpaced the evidence base: fewer than 5% of cleared U.S. Food and Drug Administration (...
John Emmett Worth, Anastasia Perez, David Wu et al.· BMJ digital health & AI· 1 citation
Cognitive safety is proposed here as a longitudinal property of the clinician-AI-organization sociotechnical system: its capacity to support or improve clinical performance without eroding independent hypothesis generation, uncertainty calibration, reasoned dissent, metacognitive control, and resilient performance when...
S. Corrao· Recenti progressi in medicin...· 0 citations
Agentic artificial intelligence (AI), comprising autonomous, goal-oriented systems capable of reasoning, planning, and coordinating multi-step clinical workflows under clinician oversight, is emerging as a transformative paradigm in modern healthcare. This study presents a systematic review of its conceptual foundation...
Multimodal artificial intelligence (AI) agents are emerging in healthcare as systems that integrate heterogeneous clinical data, foundation models (FMs), tools, and agentic workflows, but their applications and translational readiness remain unclear. We conducted a scoping review of 37 peer-reviewed studies publish...
Kai Yu, Shuang Zhou, Yu Hou et al.· npj Digital Medicine· 1 citation
This survey aims to examine the paradigm shift from predictive, assistive artificial intelligence (AI) to agentic AI in healthcare. Agentic AI refers to systems capable of perceiving, reasoning, planning and acting with adaptive autonomy, enabling more proactive and intelligent decision-making within complex health...
M. Albashrawi· Information Discovery and De...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.