Skip to content
Review Open access

RankVar: machine learning-based variant ranking and reinterpretation for rare genetic diseases.

Jul 2026 · Genome Medicine · 1 citation
Medicine

TL;DR

A machine learning algorithm called RankVar is developed to prioritize causative variants for rare diseases, based on clinical notes and genome/exome sequencing profiles, and may provide a useful framework for prioritizing variants in monogenic or oligogenic diseases.

Abstract

Background

Prior biological knowledge and phenotype information can help identify disease genes from whole genome/exome sequencing studies, but how best to incorporate external knowledge with variant data remains challenging. We developed a machine learning algorithm called RankVar to prioritize causative variants for rare diseases, based on clinical notes and genome/exome sequencing profiles.

Methods

RankVar uses a random forest classifier trained on ~ 1 million variants from the 1000 Genomes Project with spiked-in pathogenic variants. For testing, we compiled sequencing data and phenotype information from several independent datasets: 260 subjects from the Children's Hospital of Philadelphia (CHOP) with positive genetic diagnosis of various Mendelian diseases, 135 subjects from Birth Defects Biorepository (BDB), as well as 356 and 97 subjects with candidate causal variants for autism spectrum disorders from the Simons Simplex Collection (SSC) and the Simons Foundation Powering Autism Research for Knowledge (SPARK), respectively.

Results

RankVar achieves a top 10 variant accuracy of 90.0%, 81.5%, 46.1%, and 76.3% for CHOP, BDB, SSC, and SPARK, respectively, with improved performance over existing approaches. Notably, RankVar successfully identified X-linked and Y-linked disease-causal variants, such as KDM6A (p.N915Kfs5*) and SRY (p.W98X), as the top candidate variants. Moreover, we evaluated RankVar for genomic reinterpretation of 130 unsolved CHOP cases with hearing loss and successfully identified 61 candidate causal variants after manual review.

Conclusions

In summary, RankVar performed favorably relative to existing methods in our evaluation, accommodated different genetic models and X/Y chromosome variants, and may provide a useful framework for prioritizing variants in monogenic or oligogenic diseases. We anticipate that RankVar may aid in primary genetic diagnosis, genome reinterpretation of previously unsolved cases, and the discovery of novel disease genes.

Read PDF

Similar papers

#explainable ai Open access Aug 2026

aiDIVA – hybrid AI for rare disease diagnostics using evidence-based, machine learning and language models

aiDIVA is presented, an ensemble-AI combining statistical and machine learning models trained on genomic and phenotypic data to identify causal variants among tens of thousands per patient, and applies a random forest model to classify pathogenicity and generates evidence-based scores for dominant and recessive diseases.

D. Boceck, L. Laugwitz, Marc Sturm et al. · 0 citations
Review Open access Aug 2026

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

Introduction Variant interpretation remains a major bottleneck in clinical genomics, with variants of uncertain significance (VUS) representing a critical unresolved challenge due to insufficient evidence for definitive classification. Existing in silico tools exhibit variable and often inconsistent performance complicating clinical decision-making, particularly in the context of hereditary cancer genomics. Methods In this study, we developed a machine learning framework trained on 1,04,646 high-confidence ClinVar germline variants (3-star+ review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 pathogenicity scores to classify variants as Pathogenic or Benign, subsequently applying the trained model to reclassify 40894 ClinVar VUS. Train/test partitioning was performed at the variant level (80/20 split) to prevent data leakage, with hyperparameter optimization via GridSearchCV and performance assessed by 10-fold cross-validation. Four classifiers were evaluated viz. Logistic Regression, Support Vector Machine, Random Forest and XGBoost, with Random Forest achieving the highest performance (AUC-ROC = 0.9995, 95% CI: 0.9993–0.9997; 10-fold CV AUC = 0.9992 ± 0.0004). Probability thresholds of P ≥ 0.80 (Pathogenic) and P <= 0.20 (Benign) were derived from Precision-Recall curve analysis, achieving empirically validated precision of 99.63% and 99.77% respectively on held-out test variants. Results and Discussion Applied to 40,894 ClinVar VUS, the model reclassified 19393 (47.4%) as Likely Pathogenic and 8,957 (21.9%) as Likely Benign, while 12,544 (30.7%) were conservatively retained as uncertain. External validation on 7,462 ENIGMA-classified BRCA1/BRCA2 variants from the BRCA Exchange database, completely independent of the ClinVar training data demonstrated an overall concordance of 98.83% (AUC = 1.0000). Further validation of VUS reclassification against 671 variants classified as VUS in ClinVar but definitively classified by ENIGMA yielded an overall concordance of 89.57% (Pathogenic: 96.4%, Benign: 87.4%). SHAP-based explainability analysis confirmed that predictions were predominantly driven by biologically interpretable features, including CADD Phred score, VEP functional impact tier, variant consequence class and population allele frequency, consistent with ACMG/AMP evidence criteria. This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al. · 0 citations
Open access Aug 2026

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease

Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%–53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

Dagnogo Dramane, M. Treccani, L. Veschetti et al. · 0 citations
Open access Aug 2026

Machine Learning approaches for the detection of disease-causing variants in whole-genome data need to address the expression of functional genes

Gene-dosage combinations have been recognised as leading factors of disease. Given that those combinations may include dozens of genes, it is hypothesised that machine learning (ML) approaches may be useful in the classification of cases and controls and the identification of causative genes. We aimed to assess the validity of this hypothesis. Here, we have constructed a benchmark that includes real data (with ground truth knowledge) and synthetic data with known generating mechanisms and various dataset sizes and levels of noise. We trained standard statistical learning/ ML models on these datasets to classify disease phenotype. We present an analysis of how model performance varies across different synthetic genetic scenarios, and how it is impacted by dataset size. The logistic regression model was found to be the most reliable at causative gene identification across the synthetic datasets, despite not always performing the best in terms of classification performance and, in some cases, having a relatively low ROC AUC score. When our training attempts on the UK Biobank datasets failed, we performed an analysis into model performance vs dataset richness. Our results show that it is necessary to take into account the expression of functional genes in order to successfully predict disease.

Camilla Mapstone, Julia Handl, David Talavera · 0 citations
Review Open access Jul 2026

Learning Minimal Gene Programs for Disease-Aligned Representations

Identifying small, interpretable gene sets that robustly capture disease-associated variation in singlecell transcriptomic data remains a central challenge for biological interpretation and experimental followup. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-heldout evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data. Data and Code Availability All code used for data preprocessing, model training, evaluation, and feature importance analyses is available at: https://github.com/AdiVM/SLMC_single-cell. The singlecell data used in this study are publicly available human transcriptomic datasets generated by prior studies and accessible through the Gene Expression Omnibus (GEO). Analyses were performed using renal cell carcinoma single-cell RNA-seq data (GEO accession: GSE314072) and Alzheimer’s disease single-nucleus RNA-seq data from human cortex (GEO accession: GSE138852), with additional publicly available datasets used for cross-context diagnostic evaluation (GEO accessions: GSM8652069, GSE308624, and GSE227734). All datasets contain de-identified human samples and are available under standard publicuse terms via GEO. Institutional Review Board (IRB) This study analyzes de-identified, publicly available human transcriptomic data obtained from previously published studies. No new data were collected, and no identifiable private information was accessed. In accordance with institutional policy, this work was determined to constitute non-human subjects research and did not require additional IRB approval.

Adithya V. Madduri, Chirag J. Patel · 0 citations