This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over.
Abstract
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts"accountable clinical AI"from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.
Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $<$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.
Prabhjot Singh, Pritam Deka, V. Chennareddy· 0 citations
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers'safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.
Clinical AI is no longer bottlenecked only by model performance; it is bottlenecked by the accountable interaction loop through which clinicians and patients inspect evidence, test alternatives, and remain responsible for decisions under uncertainty. We argue that today's digital-twin systems, XR interfaces, and foundation-model copilots each address part of this loop, but they fail when deployed as separate products: predictions are not replayable across time, XR becomes descriptive visualization without executable state, and copilots are fluent without auditable grounding. We introduce the Metaverse Patient Digital Twin (MPDT) as a decision-grade clinical artifact defined by one requirement: every displayed claim or simulated scenario must be traceable to a versioned patient state, explicit assumptions, and replayable interaction logs. We specify minimal acceptance criteria, a reference loop in which stakeholders observe new evidence, update the twin state, run bounded simulations, generate explanations, commit decisions, and then log and monitor outcomes, along with a compact architecture that binds interoperability, simulation, governed interaction, and lifecycle controls. Finally, we outline workflows (risk stratification, diagnosis support, treatment rehearsal, training, cross-site coordination) and the evidence required to make MPDTs defensible: calibration over time, subgroup reliability, category-error prevention (observation vs simulation), and measurable workflow outcomes.
Filippo Cenacchi, Longbing Cao, Deborah Richards· Proceedings of the 32nd ACM...· 0 citations
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.
Veronica Chatrath, Bryan Zhu, George Pu et al.· 0 citations
Large Language Models (LLMs) show strong potential for clinical reasoning, yet their deployment in medical decision support is hindered by hallucinations, overconfidence, and limited transparency. We propose Dialectic Diagnosis, an agentic framework in which two heterogeneous LLM agents engage in structured argumentative interaction inspired by clinical second-opinion workflows. A Clinical Reasoner proposes candidate diagnoses, while a Skeptical Critic challenges these hypotheses by identifying omissions, cognitive biases, and unsupported reasoning. Their interaction is governed by a formal finite-state machine (FSM) enforcing a disciplined proposecritique-resolve protocol, with final decisions produced by an Arbiter agent providing calibrated confidence estimates. To ensure transparency, we introduce a Diagnostic Argument Graph that explicitly represents supporting evidence, contradictions, and missing diagnoses. Evaluations on real-world clinical datasets (MIMIC-IV and eICU) show clear gains over single-agent LLM baselines, with improved diagnostic accuracy, lower calibration error, and fewer critical diagnostic omissions. These results indicate that structured argumentative interaction between LLM agents provides a principled path toward safer and more explainable clinical AI systems.
Belkacem Chikhaoui· Annual International Compute...· 0 citations
Generative AI is entering clinical practice not as a predictor but as an actor. Contemporary healthcare deployments increasingly involve agents: language-model systems that plan over multiple steps,retrieve patient context, invoke tools, write to the electronic health record, and coordinate with otheragents. Health AI governance, however, remains overwhelmingly design-time. Premarket review,transparency labels, and reporting standards evaluate a model artifact under the assumption that behavior is a stable property of that artifact. Agentic systems violate this assumption: their effective behavioris constituted at runtime by the composition of instructions, retrieved context, tool affordances, mem
ory, and inter-agent interaction, none of which is fixed at approval time. This paper proposes ClinicalAgentOps, a framework that relocates governance from the artifact into the agent’s execution path. Itcontributes an explicit argument from the premises of design-time assurance to the necessity of inpath control, stated with its falsifying conditions; a runtime failure taxonomy for clinical agents inwhich the unit of analysis is the action trajectory rather than the input–output pair; a two-dimensionalmodel treating autonomy as a graduated, revocable, per-action grant indexed by a Clinical Action RiskTier, with a stated derivation rule from which the minimum control set for each (tier, autonomy) pairfollows; a five-plane reference architecture spanning authorization, execution, observation, assurance,and accountability; a clinical agent trace schema extending emerging generative-AI telemetry conventions with attribution, evidence, and oversight attributes, together with governance metrics computablefrom it; and a mapping from framework components to obligations under prevailing risk-management,
privacy, and medical-device regimes. This is a framework and position paper: claims about controlefficacy are advanced as falsifiable hypotheses with the study designs that would test them, not asresults
K. Bisht, R. Kumar· International Journal For Mu...· 0 citations