Skip to content
Book Open access

The Decision Twin: Metaverse Patient Digital Twins as Executable Clinical Reality

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 13157-13163 · 0 citations · 61 references

TL;DR

This work introduces the Metaverse Patient Digital Twin as a decision-grade clinical artifact defined by one requirement: every displayed claim or simulated scenario must be traceable to a versioned patient state, explicit assumptions, and replayable interaction logs.

Abstract

Clinical AI is no longer bottlenecked only by model performance; it is bottlenecked by the accountable interaction loop through which clinicians and patients inspect evidence, test alternatives, and remain responsible for decisions under uncertainty. We argue that today's digital-twin systems, XR interfaces, and foundation-model copilots each address part of this loop, but they fail when deployed as separate products: predictions are not replayable across time, XR becomes descriptive visualization without executable state, and copilots are fluent without auditable grounding. We introduce the Metaverse Patient Digital Twin (MPDT) as a decision-grade clinical artifact defined by one requirement: every displayed claim or simulated scenario must be traceable to a versioned patient state, explicit assumptions, and replayable interaction logs. We specify minimal acceptance criteria, a reference loop in which stakeholders observe new evidence, update the twin state, run bounded simulations, generate explanations, commit decisions, and then log and monitor outcomes, along with a compact architecture that binds interoperability, simulation, governed interaction, and lifecycle controls. Finally, we outline workflows (risk stratification, diagnosis support, treatment rehearsal, training, cross-site coordination) and the evidence required to make MPDTs defensible: calibration over time, subgroup reliability, category-error prevention (observation vs simulation), and measurable workflow outcomes.

Read PDF

Similar papers

Preprint Aug 2026

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over.

Augusto Bernardo Pissarra, Victor Farias DE Souza · 0 citations

EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation

Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable...

Feng-Nan Li, Heman Burre, Li-Wen Sun et al. · 0 citations
Review Open access Jul 2026

From Algorithm to Bedside: A Clinician's Framework for AI in Practice

This article proposes seven questions that clinicians can run through to evaluate any clinical AI tool in the time it takes to read an abstract, alongside a traffic-light schema for matching oversight to risk and a short list of demands clinicians should make of vendors and institutions.

Alaa Abdelqader, M. Alkhateeb, Abdullah Al-Marrawi et al. · 0 citations
#artificial intelligence Review Sep 2026

KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI

Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the propor...

Jocelyn Kang, Caroline Zhang · 0 citations
Review Open access Sep 2026

Clinical warrant in medical LLM evaluation

Medical large language models are evaluated through factuality, physician preference, source support, retrieval quality, safety, and clinical utility. These measures describe content and performance while leaving unresolved whether a generated claim meets the requirements for a specified clinical use. This article prop...

Murat Sariyar · 0 citations
#natural language process... Preprint Sep 2026

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

A methodology is implemented for evaluating large language models on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation, which shows that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lif...

Dipankar Srirag, Hao-Kai Zhao, Ashutosh Kumar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.