UshEffect-3D serves as an interpretable protein-specific framework and a supplementary ranked resource that can be used to flag unresolved variants for further clinical and experimental review.
Abstract
USH2A encodes Usherin, a 5,202-residue multi-domain protein whose missense variation is a major cause of Usher syndrome type II. Interpretation remains challenging because the protein lacks a full-length experimental structure and thousands of reported missense variants remain unresolved. In an attempt to address this gap, we developed UshEffect-3D, a gene-specific machine-learning-based variant effect predictor that integrates sequence-based evolutionary measures, structure-based biochemical properties, local interaction terms, and predicted protein stability changes to prioritise USH2A missense variants of unknown significance (VUS). To ensure that the structure-derived features were extracted from reliably modelled predicted structures, the AlphaFold2 models of 47 annotated Usherin domains were compared against experimental structures of representative domain types. The predicted structures showed strong fold concordance with representative experimental structures, with 43 domains achieving TM-scores of at least 0.5. Among binary classifiers trained using a residue position-grouped cross-validation strategy, Logistic Regression was selected as the most stable performer and achieved an MCC of 0.75, AUC of 0.96, sensitivity of 0.93 and specificity of 0.81 on a group-aware held-out test set. On a benchmarking task, UshEffect-3D performed comparably to six general-purpose variant-effect predictors evaluated on the held out test set, scoring the highest on every balanced metric but no statistically resolvable advantage over any major competitors such as PolyPhen-2, VESPAl, ESM-1b and AlphaMissense. Feature ablation displayed that sequence-derived evolutionary constraints were the sole detectable source of discrimination - withholding the conservation features reduced cross-validated MCC by 0.35 (paired 95% CI −0.45 to −0.25), whereas removing structural descriptors and predicted stability change negligibly affected model performance. Applying our model to the 2,639 ClinVar VUS, we reclassified 981 (37.2%) as likely pathogenic highlighting the extent of disease-relevant missense variation yet to be uncovered. Thus, UshEffect-3D serves as an interpretable protein-specific framework and a supplementary ranked resource that can be used to flag unresolved variants for further clinical and experimental review.
Variant effect prediction remains a key challenge in precision medicine. Computational models are increasingly successful in the characterization of missense variants in folded protein regions. However, 37% of all annotated missense variants reside in the 25% of the proteome that is intrinsically disordered, lacking positional sequence conservation and stable structures. To advance the characterization of variants in intrinsically disordered protein regions (IDRs), we combined sequence pattern searches with AlphaFold to structurally annotate 1,300 protein–protein interactions with interfaces mediated by short disordered motifs binding to folded domains in partner proteins. These interfaces were selected based on their overlap with uncertain missense variants enabling structural model-based prediction of deleterious effects of 1,187 of these variants in IDRs. Extensive experimental efforts validated the predicted interfaces and deleterious variant effects that were predicted as benign by AlphaMissense, demonstrating that the combination of sequence analysis and structural modeling can readily generate numerous testable hypotheses of variant effects on protein function in IDRs. Proteome-wide prediction and structural modeling of disordered protein interaction interfaces advance characterization of disease-associated variants in disordered protein regions.
D. Hubrich, Jesús Alvarado Valverde, C. Y. Lee et al.· Nature Structural & Molecula...· 0 citations
Purpose
As sequencing improves, identifying variants causing inherited retinal diseases (IRDs) is essential for gene therapy. Structure-based network analysis (SBNA) predicts missense variant impact based entirely on structural first principles rather than historical phenotypic or clinical outcome data, distinguishing it among contemporary missense prediction tools. Here, we expanded the application of SBNA to artificial intelligence (AI)-generated protein structures, facilitating application to all known IRD-associated proteins.
Methods
We first calculated SBNA scores for structures from the Protein Data Bank (PDB) and AI-generated structures from AlphaFold2, comparing scores for pathogenic and benign ClinVar variants. We then used these results to identify the putative genetic basis of disease for patients with IRDs, demonstrating the clinical applicability of this approach.
Results
We found a significant difference between SBNA scores for known benign and pathogenic variants across all human protein structures from the PDB (median, -0.6 vs. 1.8; P < 0.0001; AUC = 0.763) and across the corresponding AlphaFold2 structures (median, -0.2 vs. 1.9; P < 0.0001; AUC = 0.755). This difference was also significant for AlphaFold2 structures from 374 IRD-associated proteins (median, -0.4 vs. 1.9; P < 0.0001; AUC = 0.779), including 185 without available structural data. This model identified likely causative disease variants in 56% of IRD patients without a known genetic basis for disease.
Conclusions
SBNA can identify variants in human proteins that are likely to cause disease, and it can help predict variants causative of IRDs in an unbiased fashion using both AlphaFold2-generated structural models and experimental structural data.
Blake Hauser, E. Place, Yuyang Luo et al.· Investigative Ophthalmology...· 0 citations
Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.
Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al.· bioRxiv· 0 citations
Abstract Objective To determine whether computational protein‐stability predictions discriminate pathogenic from benign SCN1A missense variants, and to characterize the structural distribution of predicted destabilization among pathogenic variants. Methods On an AlphaFold3‐predicted Nav1.1 structure, FoldX, and Rosetta Cartesian ΔΔG were computed for a single ClinVar snapshot of pathogenic/likely‐pathogenic (P/LP) and benign/likely‐benign (B/LB) missense variants and its extension to ClinVar variants of uncertain significance (VUS) and gnomAD v4.1 variants; membrane‐aware RosettaMP was applied to the patch‐clamp subgroup. Pathogenic variants were stratified by functional domain. Results Pathogenic variants were more destabilizing than benign (FoldX 2.61 vs. 0.31 kcal/mol, p = 1.27 × 10−11; ROC‐AUC = 0.760), concordant with Rosetta (ROC‐AUC = 0.697; ρ = 0.660). Destabilization was domain‐dependent: pore (P‐loop/selectivity‐filter) pathogenic variants were depleted of stability‐neutral variants (0.40‐fold; Bonferroni‐adjusted p = 4.3 × 10−7), whereas S4 voltage‐sensor variants were enriched for them (2.32‐fold; p = 0.013). Across ~3300 non‐redundant variants, gnomAD‐common variants resembled benign controls and VUS were intermediate (mean ΔΔG 0.98 kcal/mol; 18.5% strongly destabilizing), with the domain pattern preserved. Among 64 patch‐clamp variants, stability did not separate gain‐ from loss‐of‐function, though gain‐of‐function variants clustered in voltage‐sensing domains and were absent from the pore. Significance Computational stability analysis thus adds a mechanistic layer complementary to the conventional gating‐dysfunction view, distinguishing a destabilized pore‐region subset—for which proteostasis impairment is a candidate, though unproven, mechanism—from a structurally tolerated S4 subset whose pathogenicity is stability‐independent. As a hypothesis‐generating rather than mechanism‐defining approach, this stratification prioritizes candidate variants—including the 18.5% of VUS that are strongly destabilizing—for direct functional and surface‐expression validation in SCN1A‐related epilepsies. Plain Language Summary We used computational modeling to predict how thousands of SCN1A genetic variants influence the stability of the Nav1.1 sodium channel protein. Disease‐causing variants tended to destabilize the protein more than benign variants, and variants in the pore region—where ions flow through the channel—were predominantly destabilizing. This is consistent with loss‐of‐function arising from misfolding and degradation of the channel protein in this subset of variants. By contrast, variants in the voltage‐sensing region were often structurally tolerated, indicating that their disease‐causing effects likely arise through a different mechanism that requires direct functional measurement to define. Accordingly, the analysis nominates a candidate pore‐region subset potentially affected by proteostasis impairment and a complementary stability‐neutral subset warranting functional evaluation.
Y. Shim, E. Kang, Naeun Kwak et al.· Epilepsia Open· 0 citations
The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.
Lars A. Eicholt, Lasse Middendorf· bioRxiv· 0 citations
Splice-altering variants cause an estimated 15-30% of genetic diseases, yet computational tools lose accuracy outside the canonical GT-AG dinucleotides, leaving intronic variants of uncertain significance (VUS) hard to interpret. Here we present MetaSplice, a 53-feature gradient-boosted ensemble integrating deep-learning splice predictions (SpliceTransformer, Pangolin), evolutionary and gene-level constraint, and splicing-regulatory motifs to score single-nucleotide variants across seven non-exonic splice regions: canonical donor and acceptor sites, donor and acceptor regions, the polypyrimidine tract (PPT), branch point, and deep intronic positions. Trained on 381,226 intronic ClinVar SNVs (35,506 pathogenic), MetaSplice achieved an area under the precision-recall curve (auPRC) of 0.995 in five-fold gene-grouped cross-validation. On a temporally held-out ClinVar test set (107,933 variants), it reached auPRC 0.987 (95% CI 0.985-0.989), outperforming SpliceAI (0.950), CADD v1.7 (0.915), SPIDEX (0.767) and S-CAP (0.253), with the largest gains in the PPT, branch-point and deep intronic regions. As a pathogenicity predictor trained on clinical significance, MetaSplice complements mechanism-specific splice-effect tools. Applied to 47,272 ClinVar splice-region VUS, it nominated 22% for functional follow-up. MetaSplice is freely available as a Docker image.
Xiaoming Liu· bioRxiv· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.