Skip to content

A visual large language foundational model for medical image recognition using clinician-oriented social media

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

A FOundational LLM Trained on ThoughtMed-1M (FOLTMed), a scalable paradigm for advancing research on clinically grounded multimodal LLMs, achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set.

Abstract

Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.

View source

Similar papers

Preprint Sep 2026

A visual large language foundational model for medical image recognition using clinician-contributed online resources

Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignm...

Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al. · 0 citations
Open access Aug 2026

Medrecord-CLIP: enhancing fundus disease diagnosis via EHR-guided vision-language pre-training

This work proposes MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations that highlights the critical value of integrating personalized clinical context to enhance the generaliza...

Lei Shi, Wenbin Zhai, Lei Yu et al. · 0 citations
Preprint Aug 2026

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.

Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al. · 0 citations
Review Open access Oct 2024

Large Language Model Benchmarks in Medical Tasks

With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including t...

L. K. Yan, Qian Niu, Ming Li et al. · 32 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Representation-guided in-context learning for medical image interpretation with multimodal large language models

Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves que...

Min-Da Zhao, Fang-Yu Hu, Yan Luo et al. · 0 citations
Open access Aug 2026

An explainable biomedical foundation model via large-scale concept-enhanced vision-language pretraining.

Artificial intelligence for medical imaging is required to be accurate and interpretable to clinicians. However, current multimodal biomedical foundation models often prioritize performance over explainability. Here we present ConceptCLIP, an explainable biomedical foundation model that achieves state-of-the-art diagno...

Yuxiang Nie, Sunan He, Yequan Bie et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.