Skip to content
#explainable ai Open access

aiDIVA – hybrid AI for rare disease diagnostics using evidence-based, machine learning and language models

Aug 2026 · npj Genomic Medicine · Vol 11 · 0 citations · 65 references
Medicine

TL;DR

aiDIVA is presented, an ensemble-AI combining statistical and machine learning models trained on genomic and phenotypic data to identify causal variants among tens of thousands per patient, and applies a random forest model to classify pathogenicity and generates evidence-based scores for dominant and recessive diseases.

Abstract

Genome sequencing enables accurate detection of genetic variants and is transforming rare disease diagnostics. While data generation is scalable, prioritization and clinical interpretation remain challenging, often requiring expert manual classification. AI-driven decision support systems are therefore needed to assist in causal variant identification or to fully automate large-scale re-analysis of unsolved cases. Existing tools often estimate variant impact on protein function, but few integrate genomic, phenotypic, and clinical annotation data for diagnosis. We present aiDIVA, an ensemble-AI combining statistical and machine learning models trained on genomic and phenotypic data to identify causal variants among tens of thousands per patient. aiDIVA applies a random forest model to classify pathogenicity and generates evidence-based scores for dominant and recessive diseases. These predictions are integrated with clinical metadata to prioritize the most likely causal variants. Large language models further refine and explain results. The aiDIVA-meta model consolidates all scores into a ranked list. aiDIVA-meta reported the causal variant among the top-3 candidates in 97.4% of a pre-training collected cohort with prior evidence in ClinVar or HGMD, and in 93.3% of a post-training collected cohort of previously unreported variants.

Read PDF

Similar papers

Review Open access Jul 2026

RankVar: machine learning-based variant ranking and reinterpretation for rare genetic diseases.

A machine learning algorithm called RankVar is developed to prioritize causative variants for rare diseases, based on clinical notes and genome/exome sequencing profiles, and may provide a useful framework for prioritizing variants in monogenic or oligogenic diseases.

Yuan Zhang, Mian Umair Ahsan, Peng Wang et al. · 1 citation
Review Open access Jul 2026

Artificial Intelligence and Genomic Data Analysis: New Frontiers in Precision Medicine

A clinically oriented, pipeline-based synthesis of contemporary AI applications in genomic medicine, focusing on factors that determine model robustness and clinical utility, and common sources of failure in real-world genomic AI systems.

Alexandra-Maria Blaga, Răzvan-Octavian Mihuț, A. Treteanu et al. · 0 citations

Explainable AI for analyzing cancer outcomes using large-scale genome sequencing data

Metastatic cancer remains a leading cause of global mortality, yet accurate prognosis is frequently hampered by high-dimensional molecular features and heterogeneous clinical presentations. While traditional staging systems and linear models provide a foundational risk assessment, they often fail to capture the complex, nonlinear interactions between metastatic topology, genomic burden, and functional sequence variation. To address this, recent advances in machine learning and genomic foundation models present a transformative opportunity to integrate diverse data types into an explainable predictive framework. Consequently, this research developed a multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates. Additionally, the framework aimed to surface sequence-level disease drivers by implementing joint variant calling from RNA-seq data and leveraging transformer-based architectures. The study employed a two-track methodological approach encompassing populationscale modeling and sequence-level deep learning. For the population-scale aim, a retrospective analysis was conducted on the Memorial Sloan Kettering-Metastatic cohort, consisting of 25,775 patients. Five distinct classifiers XGBoost, Logistic Regression, Random Forest, Decision Tree, and Naive Bayes were trained on a balanced subset of 20,338 patients utilizing an 80/20 stratified split. Model explainability was established through Shapley Additive Explanations (SHAP), while survival dynamics were evaluated using Kaplan-Meier estimates, Cox proportional hazards models, and an XGBoost-Cox variant. Concurrently, a pilot study involving 60 individuals, comprising 30 breast cancer cases and 30 controls, investigated sequence-level drivers using RNAseq data. A joint variant calling pipeline generated a unified genomic variant call format for association testing, and three genomic foundation models DNABERT-2, HyenaDNA, and Nucleotide Transformer were fine-tuned for 50 epochs on variantcentered windows spanning 100 base pairs in either direction to classify case versus control status. The results revealed stark contrasts in performance between the clinical and genomic modeling tracks. In survivability predictions, XGBoost emerged as the superior classifier, achieving an accuracy of 0.74 and an AUC of 0.82, while the XGBoost-Cox model outperformed the traditional Cox model with a C-index of 0.70 compared to 0.66. Through explainability and hazard-based analyses, metastatic site count, tumor mutational burden, the fraction of the genome altered, and the presence of liver and bone metastases were identified as the most potent prognostic indicators across pan-cancer and cancer-specific models. Conversely, the sequence-level transformer models exhibited severe overfitting, with test performance remaining near stochastic levels between 49 percent and 51 percent accuracy. Although DNABERT-2 achieved the highest nominal accuracy at 50.63 percent and HyenaDNA showed superior computational efficiency, the pilot ultimately indicated that fine-tuning transformers on raw sequences in small cohorts is heavily limited by a high signal-to-noise ratio and the polygenic complexity of cancer. Ultimately, this research demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards. However, future sequence-level deep learning efforts must pivot toward using frozen transformer embEd. D.ings or larger, multi-center cohorts to ensure equitable and generalizable clinical adoption.

P. Nalela · 0 citations
Open access Aug 2026

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

Karthik V, S. Prejesh, Sumedh Deepak Kudale et al. · 0 citations
Review Open access Aug 2026

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

Introduction Variant interpretation remains a major bottleneck in clinical genomics, with variants of uncertain significance (VUS) representing a critical unresolved challenge due to insufficient evidence for definitive classification. Existing in silico tools exhibit variable and often inconsistent performance complicating clinical decision-making, particularly in the context of hereditary cancer genomics. Methods In this study, we developed a machine learning framework trained on 1,04,646 high-confidence ClinVar germline variants (3-star+ review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 pathogenicity scores to classify variants as Pathogenic or Benign, subsequently applying the trained model to reclassify 40894 ClinVar VUS. Train/test partitioning was performed at the variant level (80/20 split) to prevent data leakage, with hyperparameter optimization via GridSearchCV and performance assessed by 10-fold cross-validation. Four classifiers were evaluated viz. Logistic Regression, Support Vector Machine, Random Forest and XGBoost, with Random Forest achieving the highest performance (AUC-ROC = 0.9995, 95% CI: 0.9993–0.9997; 10-fold CV AUC = 0.9992 ± 0.0004). Probability thresholds of P ≥ 0.80 (Pathogenic) and P <= 0.20 (Benign) were derived from Precision-Recall curve analysis, achieving empirically validated precision of 99.63% and 99.77% respectively on held-out test variants. Results and Discussion Applied to 40,894 ClinVar VUS, the model reclassified 19393 (47.4%) as Likely Pathogenic and 8,957 (21.9%) as Likely Benign, while 12,544 (30.7%) were conservatively retained as uncertain. External validation on 7,462 ENIGMA-classified BRCA1/BRCA2 variants from the BRCA Exchange database, completely independent of the ClinVar training data demonstrated an overall concordance of 98.83% (AUC = 1.0000). Further validation of VUS reclassification against 671 variants classified as VUS in ClinVar but definitively classified by ENIGMA yielded an overall concordance of 89.57% (Pathogenic: 96.4%, Benign: 87.4%). SHAP-based explainability analysis confirmed that predictions were predominantly driven by biologically interpretable features, including CADD Phred score, VEP functional impact tier, variant consequence class and population allele frequency, consistent with ACMG/AMP evidence criteria. This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al. · 0 citations

Related blog posts