Skip to content
#protein folding Review Open access

Determinants of missense pathogenicity in Usherin: Evaluating sequence and predicted-structure features

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

UshEffect-3D serves as an interpretable protein-specific framework and a supplementary ranked resource that can be used to flag unresolved variants for further clinical and experimental review.

Abstract

USH2A encodes Usherin, a 5,202-residue multi-domain protein whose missense variation is a major cause of Usher syndrome type II. Interpretation remains challenging because the protein lacks a full-length experimental structure and thousands of reported missense variants remain unresolved. In an attempt to address this gap, we developed UshEffect-3D, a gene-specific machine-learning-based variant effect predictor that integrates sequence-based evolutionary measures, structure-based biochemical properties, local interaction terms, and predicted protein stability changes to prioritise USH2A missense variants of unknown significance (VUS). To ensure that the structure-derived features were extracted from reliably modelled predicted structures, the AlphaFold2 models of 47 annotated Usherin domains were compared against experimental structures of representative domain types. The predicted structures showed strong fold concordance with representative experimental structures, with 43 domains achieving TM-scores of at least 0.5. Among binary classifiers trained using a residue position-grouped cross-validation strategy, Logistic Regression was selected as the most stable performer and achieved an MCC of 0.75, AUC of 0.96, sensitivity of 0.93 and specificity of 0.81 on a group-aware held-out test set. On a benchmarking task, UshEffect-3D performed comparably to six general-purpose variant-effect predictors evaluated on the held out test set, scoring the highest on every balanced metric but no statistically resolvable advantage over any major competitors such as PolyPhen-2, VESPAl, ESM-1b and AlphaMissense. Feature ablation displayed that sequence-derived evolutionary constraints were the sole detectable source of discrimination - withholding the conservation features reduced cross-validated MCC by 0.35 (paired 95% CI −0.45 to −0.25), whereas removing structural descriptors and predicted stability change negligibly affected model performance. Applying our model to the 2,639 ClinVar VUS, we reclassified 981 (37.2%) as likely pathogenic highlighting the extent of disease-relevant missense variation yet to be uncovered. Thus, UshEffect-3D serves as an interpretable protein-specific framework and a supplementary ranked resource that can be used to flag unresolved variants for further clinical and experimental review.

Read PDF

Similar papers

Open access Jul 2026

Variant characterization in the intrinsically disordered human proteome

Variant effect prediction remains a key challenge in precision medicine. Computational models are increasingly successful in the characterization of missense variants in folded protein regions. However, 37% of all annotated missense variants reside in the 25% of the proteome that is intrinsically disordered, lacking positional sequence conservation and stable structures. To advance the characterization of variants in intrinsically disordered protein regions (IDRs), we combined sequence pattern searches with AlphaFold to structurally annotate 1,300 protein–protein interactions with interfaces mediated by short disordered motifs binding to folded domains in partner proteins. These interfaces were selected based on their overlap with uncertain missense variants enabling structural model-based prediction of deleterious effects of 1,187 of these variants in IDRs. Extensive experimental efforts validated the predicted interfaces and deleterious variant effects that were predicted as benign by AlphaMissense, demonstrating that the combination of sequence analysis and structural modeling can readily generate numerous testable hypotheses of variant effects on protein function in IDRs. Proteome-wide prediction and structural modeling of disordered protein interaction interfaces advance characterization of disease-associated variants in disordered protein regions.

D. Hubrich, Jesús Alvarado Valverde, C. Y. Lee et al. · 0 citations
Open access Aug 2026

Structure-Based Network Analysis of AlphaFold Structure Predictions Identifies Putative Causative Variants of Inherited Retinal Disease.

Purpose As sequencing improves, identifying variants causing inherited retinal diseases (IRDs) is essential for gene therapy. Structure-based network analysis (SBNA) predicts missense variant impact based entirely on structural first principles rather than historical phenotypic or clinical outcome data, distinguishing it among contemporary missense prediction tools. Here, we expanded the application of SBNA to artificial intelligence (AI)-generated protein structures, facilitating application to all known IRD-associated proteins. Methods We first calculated SBNA scores for structures from the Protein Data Bank (PDB) and AI-generated structures from AlphaFold2, comparing scores for pathogenic and benign ClinVar variants. We then used these results to identify the putative genetic basis of disease for patients with IRDs, demonstrating the clinical applicability of this approach. Results We found a significant difference between SBNA scores for known benign and pathogenic variants across all human protein structures from the PDB (median, -0.6 vs. 1.8; P < 0.0001; AUC = 0.763) and across the corresponding AlphaFold2 structures (median, -0.2 vs. 1.9; P < 0.0001; AUC = 0.755). This difference was also significant for AlphaFold2 structures from 374 IRD-associated proteins (median, -0.4 vs. 1.9; P < 0.0001; AUC = 0.779), including 185 without available structural data. This model identified likely causative disease variants in 56% of IRD patients without a known genetic basis for disease. Conclusions SBNA can identify variants in human proteins that are likely to cause disease, and it can help predict variants causative of IRDs in an unbiased fashion using both AlphaFold2-generated structural models and experimental structural data.

Blake Hauser, E. Place, Yuyang Luo et al. · 0 citations
Open access Aug 2026

Large-scale structure prediction of DUF-containing protein-protein interactions

Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.

Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al. · 0 citations
Open access Jul 2026

Computational protein stability analysis of SCN1A missense variants reveals domain‐dependent stability patterns

Abstract Objective To determine whether computational protein‐stability predictions discriminate pathogenic from benign SCN1A missense variants, and to characterize the structural distribution of predicted destabilization among pathogenic variants. Methods On an AlphaFold3‐predicted Nav1.1 structure, FoldX, and Rosetta Cartesian ΔΔG were computed for a single ClinVar snapshot of pathogenic/likely‐pathogenic (P/LP) and benign/likely‐benign (B/LB) missense variants and its extension to ClinVar variants of uncertain significance (VUS) and gnomAD v4.1 variants; membrane‐aware RosettaMP was applied to the patch‐clamp subgroup. Pathogenic variants were stratified by functional domain. Results Pathogenic variants were more destabilizing than benign (FoldX 2.61 vs. 0.31 kcal/mol, p = 1.27 × 10−11; ROC‐AUC = 0.760), concordant with Rosetta (ROC‐AUC = 0.697; ρ = 0.660). Destabilization was domain‐dependent: pore (P‐loop/selectivity‐filter) pathogenic variants were depleted of stability‐neutral variants (0.40‐fold; Bonferroni‐adjusted p = 4.3 × 10−7), whereas S4 voltage‐sensor variants were enriched for them (2.32‐fold; p = 0.013). Across ~3300 non‐redundant variants, gnomAD‐common variants resembled benign controls and VUS were intermediate (mean ΔΔG 0.98 kcal/mol; 18.5% strongly destabilizing), with the domain pattern preserved. Among 64 patch‐clamp variants, stability did not separate gain‐ from loss‐of‐function, though gain‐of‐function variants clustered in voltage‐sensing domains and were absent from the pore. Significance Computational stability analysis thus adds a mechanistic layer complementary to the conventional gating‐dysfunction view, distinguishing a destabilized pore‐region subset—for which proteostasis impairment is a candidate, though unproven, mechanism—from a structurally tolerated S4 subset whose pathogenicity is stability‐independent. As a hypothesis‐generating rather than mechanism‐defining approach, this stratification prioritizes candidate variants—including the 18.5% of VUS that are strongly destabilizing—for direct functional and surface‐expression validation in SCN1A‐related epilepsies. Plain Language Summary We used computational modeling to predict how thousands of SCN1A genetic variants influence the stability of the Nav1.1 sodium channel protein. Disease‐causing variants tended to destabilize the protein more than benign variants, and variants in the pore region—where ions flow through the channel—were predominantly destabilizing. This is consistent with loss‐of‐function arising from misfolding and degradation of the channel protein in this subset of variants. By contrast, variants in the voltage‐sensing region were often structurally tolerated, indicating that their disease‐causing effects likely arise through a different mechanism that requires direct functional measurement to define. Accordingly, the analysis nominates a candidate pore‐region subset potentially affected by proteostasis impairment and a complementary stability‐neutral subset warranting functional evaluation.

Y. Shim, E. Kang, Naeun Kwak et al. · 0 citations
Open access Aug 2026

A discrete protein subset drives structure prediction discordance in orphan proteins

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Lars A. Eicholt, Lasse Middendorf · 0 citations
Open access Jul 2026

MetaSplice: an ensemble pathogenicity predictor for intronic splice variants

Splice-altering variants cause an estimated 15-30% of genetic diseases, yet computational tools lose accuracy outside the canonical GT-AG dinucleotides, leaving intronic variants of uncertain significance (VUS) hard to interpret. Here we present MetaSplice, a 53-feature gradient-boosted ensemble integrating deep-learning splice predictions (SpliceTransformer, Pangolin), evolutionary and gene-level constraint, and splicing-regulatory motifs to score single-nucleotide variants across seven non-exonic splice regions: canonical donor and acceptor sites, donor and acceptor regions, the polypyrimidine tract (PPT), branch point, and deep intronic positions. Trained on 381,226 intronic ClinVar SNVs (35,506 pathogenic), MetaSplice achieved an area under the precision-recall curve (auPRC) of 0.995 in five-fold gene-grouped cross-validation. On a temporally held-out ClinVar test set (107,933 variants), it reached auPRC 0.987 (95% CI 0.985-0.989), outperforming SpliceAI (0.950), CADD v1.7 (0.915), SPIDEX (0.767) and S-CAP (0.253), with the largest gains in the PPT, branch-point and deep intronic regions. As a pathogenicity predictor trained on clinical significance, MetaSplice complements mechanism-specific splice-effect tools. Applied to 47,272 ClinVar splice-region VUS, it nominated 22% for functional follow-up. MetaSplice is freely available as a Docker image.

Xiaoming Liu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.