Jul 2026· Journal of Vascular Diseases· Vol 5, pp. 29· 0 citations· 46 references
TL;DR
A Large Language Model-guided retrieval-aware framework for generating realistic synthetic EHR data on a large scale to train ML models to predict cardiovascular risks accurately and achieves high performance in predicting CVD.
Abstract
Background: Increasing access to Electronic Health Records (EHRs) has enabled the development of Machine Learning (ML) models to predict early cardiovascular disease (CVD) risk. Nevertheless, real EHR data is usually limited in availability and distribution because of confidentiality issues, legal limitations, and insufficient access. Synthetic data generation has become a promising approach to overcome such difficulties. Objectives: This study proposes a Large Language Model (LLM)-guided retrieval-aware framework for generating realistic synthetic EHR data on a large scale to train ML models to predict cardiovascular risks accurately. Methods: The framework uses a Tabular Denoising Diffusion Probabilistic Model to learn the underlying distribution of the original dataset and generate an initial synthetic dataset. To improve the clinical plausibility of the generated data, a knowledge-guided refinement module with LLaMA 2 13B combined with Retrieval-Augmented Generation (RAG) is introduced. The LLM analyzes statistical trends in real and synthetic data while retrieving relevant medical information to detect and correct clinically implausible correlations. Results: Experimental findings show that the optimized synthetic dataset preserves important statistical features of the original data and achieves high performance in predicting CVD. Conclusions: Thus, the framework offers a scalable, privacy-conservative method for generating realistic synthetic healthcare data suitable for medical research.
Cardiovascular diseases (CVDs) are among the leading causes of mortality worldwide. Early detection and risk stratification are critical for preventive care. Traditional machine learning (ML) models can predict heart disease risk but often lack interpretability and fail to integrate with real-time clinical data. Recent advances in fine-tuned large language models (Custom GPT) offer natural language explanations but are limited by insufficient interoperability with heterogeneous healthcare data sources. This study aims to design and evaluate a Model Context Protocol (MCP)–enabled Custom GPT framework that integrates ML-based predictive models with external healthcare systems—including EHRs, laboratory APIs, and wearable devices—to deliver context-aware, explainable, and clinically actionable heart disease risk predictions. The experimental evaluation was conducted on a validated cardiovascular dataset containing 303 patient records and 14 clinically relevant attributes derived from publicly available clinical repositories. Experimental evaluation demonstrated improved predictive accuracy (approximately 88% with the XGBoost ensemble) and robustness compared to standalone models. MCP integration enabled dynamic contextual awareness, reduced latency in tool orchestration, and enriched interpretability through RAG-based explanations. Clinician and patient evaluations confirmed enhanced usability and transparency. This approach paves the way for broader adoption of agentic AI in clinical workflows.
Neha Gupta, Bhawna Singla· Journal of Electronic &...· 0 citations
Diabetes is a major risk factor for the development of cardiovascular issues which contribute to cardiovascular disease (CVD) being a leading cause of mortality worldwide. However, traditional machine learning methods are not widely adopted in healthcare systems because they lack interpretability, which is important for early and accurate CVD risk prediction and for ruling out effective clinical intervention. In this research, a hybrid architecture is proposed that incorporates diabetes related datasets as well as explainable artificial intelligence (XAI) methodologies that could improve the prediction power and transparency of the models. The proposed approach combines different datasets at the level of features and includes rigorous data pre-processing to detect metabolic and cardiovascular risk factors. Some of the significant clinical parameters are age, BMI, glucose, cholesterol, and blood pressure. These are standardized to create a single dataset which may be utilized for predictive modelling. The employment of two XAI approaches, SHAP (SHapley Additive Explanations) with tree-based ensemble models and integrated gradients with transformer based topologies, ensures both performance and interpretability. The technique improves confidence and usefulness in clinical settings by offering accurate predictions and explanations for the model’s judgments that are relevant to the circumstance. It is also utilized for visual investigation of clinical correlations of diabetes and cardiovascular disease and identify crucial risk variables. The results suggest that merging explainability approaches with powerful machine learning can considerably boost early identification and risk assessment. The proposed approach contributes to enhanced healthcare decision-making, offering a scalable, interpretable and dependable solution for cardiovascular disease prediction.
K. Deepthi, P. Bhargavi· International journal of com...· 0 citations
The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data and demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.
S. Bin Akter, S. Akter, D. Eisenberg et al.· medRxiv· 0 citations
Motivation: Rare disease (RD) diagnosis is frequently delayed due to the similarities in symptoms to common disease variants. Machine Learning Algorithms applied to Electronic Health Records show promise for accelerating the diagnosis; however, legal and privacy concerns pose significant barriers. To address these issues, Synthetic Data Generation is an alternative method for obtaining Electronic Health Records and can be applied with any Machine Learning algorithm for benchmarking and development purposes. Despite the availability of Synthetic Data Generation algorithms, support for generating a subset of patients that differ in a definable degree from the majority to simulate patients with RD is often lacking. Results: We present SYNRARE, a graphical user interface based on the Synthea framework that enables easier modification and generation of synthetic Electronic Health Records of RD patients, which differ only to a definable degree from patients with common diseases, thereby enabling the benchmarking and testing of algorithms under controlled technical conditions. SYNRARE enables researchers to rapidly benchmark their Machine Learning algorithms across any scenario. Availability and implementation: SYNRARE, including detailed instructions for installing, is available at https://gitlab.sdu.dk/screen4care/synrare.
Nicolai Dinh Khang Truong, Richard Rottger· 0 citations
This approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships and successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support system.
U. Luke, P. Asuquo, Victor Anaga et al.· E3S Web of Conferences· 0 citations