A scalable, knowledge graph driven big data framework for explainable clinical decision support that unifies heterogeneous healthcare data into a semantically structured representation and provides transparent, traceable decision paths through knowledge graph reasoning addressing key challenges of interpretability and trust in clinical AI systems is proposed.
Abstract
The increasing volume and heterogeneity of healthcare data pose significant challenges for developing reliable and interpretable clinical decision support systems. Conventional machine learning approaches often struggle to integrate structured electronic health records, real-time patient inputs, and unstructured clinical narratives at scale, limiting their effectiveness in complex medical settings. This study proposes a scalable, knowledge graph driven big data framework for explainable clinical decision support that unifies heterogeneous healthcare data into a semantically structured representation. The framework integrates RDF-based semantic modeling, domain-specific natural language processing for entity extraction, and graph-based reasoning to map patient-reported symptoms to evidence-based treatment guidelines. Large-scale clinical data from the MIMIC-III database comprising over 40,000 hospital admissions real-time patient records, and international clinical protocols from the International Diabetes Federation (IDF) are incorporated to enable dynamic, data-driven decision making. Experimental evaluation demonstrates strong predictive performance in detecting critical diabetic conditions under controlled settings, achieving perfect precision and recall for hypoglycemia and a recall of 0.90 for diabetic ketoacidosis. A macro-averaged F1-score of 0.79 is achieved across all four diabetic condition classes, comparing favorably with rule-based clinical decision support systems and classical machine learning baselines while offering superior explainability. In addition to predictive accuracy, the framework provides transparent, traceable decision paths through knowledge graph reasoning addressing key challenges of interpretability and trust in clinical AI systems. The results highlight the effectiveness of knowledge graph based big data integration for scalable, explainable, and guideline-compliant clinical decision support. The proposed framework is generalizable to other data-intensive healthcare applications, offering a robust foundation for next-generation big data analytics and intelligent decision systems.
Clinical decision-making is frequently hindered by fragmented electronic health records and limited interoperability among heterogeneous healthcare information systems. This article presents a step-by-step protocol for implementing an interoperable web platform that integrates clinical data via the Fast Healthcare Interoperability Resources (FHIR) standard, along with Retrieval-Augmented Generation (RAG), Large Language Models (LLMs), and a multi-agent clinical reasoning framework to support medical data analysis and clinical decision support. The protocol describes the complete workflow, including computational environment configuration, clinical dataset preprocessing, FHIR-based data integration, vector database construction, retrieval configuration, prompt engineering, multi-agent orchestration, and system evaluation. Representative results demonstrate the platform's ability to generate clinically relevant and contextually consistent responses while improving semantic interoperability across heterogeneous data sources. System performance was evaluated using complementary quantitative and semantic metrics, including BLEU, ROUGE, BERTScore, and cosine similarity. A qualitative assessment was conducted using publicly available, fully anonymized benchmark healthcare datasets to evaluate the reproducibility of the proposed methodological workflow. The proposed architecture combines standardized healthcare interoperability with retrieval-enhanced language models to improve contextual reasoning, reduce hallucination, and support reproducible AI-assisted clinical workflows. This protocol provides a scalable and reproducible framework for researchers and developers seeking to implement interoperable, privacy-aware, and intelligent healthcare systems for clinical decision support, medical data analysis, and future translational research.
Marcello Carvalho Dos Reis, Rafaelly Rios Dos Santos, Daniel Santos da Silva et al.· Journal of Visualized Experi...· 0 citations
Initial experiments on heart-failure-focused clinical question answering show that CGX improves evidence retrieval quality and perceived answer reliability over conventional retrieval methods, while reducing total graph construction time by 69.7% under the same input corpus and hardware setting.
Dat Nguyen, Anh N. Le, Binh T. D. Trinh et al.· Journal of Biomedical Inform...· 0 citations
This approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships and successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support system.
U. Luke, P. Asuquo, Victor Anaga et al.· E3S Web of Conferences· 0 citations
A clinical knowledge graph construction and refinement framework that combines multi-agent relation proposal, ontology-aware normalization, deterministic evidence scoring, and JEPA-based latent refinement is proposed.
Kushagra Yadav, N. Prabhath, Amit Lamba et al.· 0 citations
This study frames cancer classification as a complex decision-support problem rather than a purely medical task. Drawing on theories of complex adaptive systems (CAS) and data-driven governance, we examine how machine-learning-based decision support systems (DSS) transform large-scale, multi-omics healthcare data into actionable risk stratification outputs. Public datasets from TCGA, CGGA, GEO, and the World Health Organization are used as a case context to evaluate algorithmic decision models, including Random Forest, Support Vector Machines, XGBoost, and SHAP-based explainability tools. Results show that these models achieve stable performance (accuracy 85-92%, AUC 0.88-0.95) while maintaining interpretability and decision reliability. By comparing healthcare analytics with tourism management and smart city governance, the study demonstrates that the same DSS logic applies across complex systems characterized by uncertainty, multi-source data, and nonlinear outcomes. The findings highlight the transferability of machine-learning-driven decision support methodologies to tourism destination management, urban governance, and other data-intensive policy domains.
Shengyu Gu· Journal of Computer Science...· 0 citations
Access to holistic, multimodal data improves the performance of Artificial Intelligence (AI) in medical classification tasks compared to utilizing single modalities or data sources. However, the inherent heterogeneity and complexity of clinical real-world data pose significant challenges to structured data analysis and AI application. This heterogeneity includes missing values, multiple time points, diverse modalities, and inconsistent formats and semantics. Data harmonization prior to data integration tackles this challenge but remains resource-intensive and error-prone, limiting the scalability and reproducibility of holistic, AI-driven decision support on clinical real-world data. We therefore propose PatTree, a graph-based, holistic representation of patients that can be derived from real-world clinical data through the automated structuring of multimodal clinical data. PatTree enables early-stage data integration without relying on pre-standardized inputs. While representing heterogeneous clinical data within a unified knowledge graph, PatTree preserves the semantic relationships between data elements across modalities and data sources, facilitating interoperability and machine-interpretable data access. Using a subset of the ADNI-1 cohort (n = 763), we demonstrate that classification of patients is directly feasible on PatTree reaching state-of-the-art classification performance. In the three-class classification task distinguishing Alzheimer's disease, mild cognitive impairment, and cognitively normal individuals, we achieve a balanced accuracy of 98.5% and an F$_1$ score of 0.987 on the held-out test set. Our results show that assumption-free, automated structuring of multimodal medical data can serve as a scalable foundation for clinical AI pipelines bypassing tedious data preparation and standardization.
Julia Gehrmann, Lars Quakulinski, H. Naseem et al.· 0 citations