Agentic Foundation Models for Explainable Multimodal Clinical Intelligence: A Human-Centered Framework for Early Disease Diagnosis and Personalized Treatment Planning
Aug 2026· Natural Resources for Human Health· 0 citations
TL;DR
A framework that treats diagnosis as a coordination problem rather than a modeling one is described and improved on both across accuracy, F1-score, explanation faithfulness and clinician-rated trust, at an added latency of roughly five seconds per case.
Abstract
Diagnostic work rarely rests on one kind of evidence. A clinician assembles the history, the imaging, the laboratory panel and, increasingly, genomic results, and does so under time pressure. Large language models perform well on each of these in isolation, yet most deployed systems still reason over a single modality, produce explanations after the fact rather than from the evidence actually used, and give the clinician no route to contest an output. This paper describes a framework that treats diagnosis as a coordination problem rather than a modeling one. Perception agents handle text, imaging and structured data separately; an orchestrator decomposes the task and routes sub-tasks among specialist agents; and a distinct explanation agent is queried after each recommendation to retrieve the inputs that drove it. Because that agent runs independently of the diagnostic pathway, its output is less likely to be a rationalization of an answer already reached. We evaluate on a pilot cohort of three cases drawn from the dataset comparing against a zero-shot baseline and a text-only specialist model. The framework improved on both across accuracy, F1-score, explanation faithfulness and clinician-rated trust, at an added latency of roughly five seconds per case. These are feasibility results from a small cohort, and we set out the prospective validation that would be needed before any claim of clinical readiness.
It is argued that each of the four challenges facing responsible development is addressed by a clinical testing harness: a structured evaluation environment comprising scenario libraries built from clinical edge cases, full-trajectory observability, explicit escalation testing, and staged evidence thresholds tied to sc...
E. Waisberg, Joseph W. Guarnieri· Annals of Biomedical Enginee...· 0 citations
Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evide...
Min-Ye Shao, Chao-Hui Yu, Yi-Xuan Wu et al.· 0 citations
Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format does not reflect how clinical diagnosis is performed in practice. As a result, they remain limited in complex diagnostic settings. We developed Debate-Mixture-of-Agents (DMoA), a novel multi-agent framework that...
Changda Xia, L. Ouyang, Hui-Min Wang et al.· 0 citations
General-purpose language models generate fluent health reports that can fabricate derived clinical metrics. In an illustrative comparison on identical two-week CGM and meal data, leading foundation models produced reports with invented MAGE values, inflated meal counts, and unreferenced complication-risk projections: f...
A. Diament, G. Sapir, M. Gorodetski et al.· medRxiv· 0 citations
Medical Structured Multimodal Memory (MSM-Mem), an agentic memory framework that enables medical AI agents to evolve through accumulated clinical experiences, is proposed, offering a viable pathway toward medical AI agents capable of evolving their reasoning competence in a manner analogous to the way clinicians learn...
Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these systems becomes a serious design concern. This is espec...
Del Coburn, Scott Sanner, Daniel A. Silver· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.