Skip to content

Deep mutational scanning of CYP2C9, CYP2C19, and NUDT15 shows that pharmacogene variant interpretation requires assay-specific functional data.

Aug 2026 · G3 · 0 citations
Medicine

TL;DR

This dimensionality dominates the data, and a supervised ESM-2 sequence baseline was benchmarked against the ESM1v zero-shot ensemble and AlphaMissense under position-based 5-fold cross-validation, together with three architectural extensions: AlphaFold structural features, multi-task learning across paired assays, and contact-graph neural networks.

Abstract

Pharmacogene missense variants can disrupt protein stability, catalytic competence, or substrate handling through distinct mechanisms. General-purpose predictors estimate clinical pathogenicity as a single scalar, whereas pharmacogene interpretation requires knowing which biochemical dimension a variant perturbs, since that determines whether reduced function is substrate-dependent. Five deep mutational scanning datasets comprising 26,198 missense variants across CYP2C9, CYP2C19, and NUDT15 were assembled from MaveDB. Paired assays showed that this dimensionality dominates the data: 28% of CYP2C9 variants (1,236 of 4,421) decoupled catalytic activity from abundance, and 48% of NUDT15 variants (1,364 of 2,844) decoupled thiopurine sensitivity from stability, with CYP2C9 discordance concentrating at substrate-channel residues. AlphaMissense, a representative general-purpose pathogenicity predictor, scored these classes in line with its clinical training objective rather than the assayed biochemistry, assigning likely-benign scores to 38 of 195 stable-but-dead CYP2C9 variants and likely-pathogenic scores to 140 of 222 destabilized but thiopurine-resistant NUDT15 variants. To test whether this dimensionality is recoverable, a supervised ESM-2 sequence baseline was benchmarked against the ESM1v zero-shot ensemble and AlphaMissense under position-based 5-fold cross-validation, together with three architectural extensions: AlphaFold structural features, multi-task learning across paired assays, and contact-graph neural networks. The baseline reached Pearson r of 0.54-0.72, matching or marginally exceeding both comparators, and no extension improved upon it. Trained directly on each assay, it nonetheless recovered the paired-assay difference at r = 0.28 for CYP2C9 and 0.43 for NUDT15, separating discordant variants at AUROC 0.60 and 0.51. Pharmacogene interpretation therefore requires assay-specific, substrate-aware functional measurements rather than a single generic score.

Read PDF

Similar papers

Open access Aug 2026

Homology-Based Variant-Effect Predictors Break Down on Cytochrome P450 Pharmacogenes

Cytochrome P450 (CYP) enzymes metabolize roughly three-quarters of clinically used drugs; genetic variation in these enzymes is a leading source of interindividual differences in drug response. Predicting a variant’s functional effect is therefore critical, yet the consequences of most CYP variants remain unknown. Many state-of-the-art variant-effect predictors rest on a homology-based paradigm that scores variants by evolutionary conservation — an assumption that pharmacogenes including CYPs violate. Indeed, focusing on human CYPs, we show that homology-based models fail systematically. AlphaMissense (AM) assigns variants to its “ambiguous” class at nearly twice the proteome-wide rate across six CYPs, and within that class the scores are essentially uncorrelated with CYP2C9 DMS activity (Spearman’s ρ = 0.069); Evolutionary Scale Modeling 2 (ESM-2) shows the same pattern. We further hypothesized that non-homology-based features (sequence position, substitution chemistry, binding-site distance, and secondary structure) might help resolve the ambiguous calls, but their explanatory power is weak. Using CYP2C9 DMS activity as ground truth, we built a k-nearest-neighbors model over ESM-2 embeddings and ensembled it with AM and ESM-2 masked marginal probability, improving the ambiguous-class correlation roughly ten-fold, from ρ = 0.069 to 0.715 (overall ρ from 0.638 to 0.825). However, this markedly improved accuracy does not translate into agreement with clinical annotations. Drawing on evidence that a variant’s effect can depend on the drug, we hypothesize that substrate identity is the key missing feature in current models, and that predicting function for these multi-substrate enzymes may require redefining function as substrate-conditioned.

Helen Z. Xu, I. Samori, Gowri Nayar et al. · 0 citations
Open access Aug 2026

Assessing Computational Models for Pharmacogenomic Variant Interpretation

Accurately predicting the effects of pharmacogenomic variants is essential for the development of personalized therapeutic strategies, as genetic variability can influence drug response differently across patients. Here, we assessed several computational approaches using a dataset of pharmacogenomic variants with either clinical annotations or functional characterization by deep mutational scanning, compiled from the literature, with an additional focus on CYP2C9, a clinically relevant drug-metabolizing enzyme. Our results show that, despite recent methodological advances, substantial room for improvement remains. In particular, current methods struggle to distinguish gain-of-function variants associated with increased drug clearance and fast-metabolizer phenotypes from neutral variants, whereas loss-of-function variants that reduce drug clearance are predicted more accurately. The integration of structural and evolutionary information appears to be a key strategy for improving performance, with the coevolution-based StructureDCA method achieving the highest accuracy compared with classical genetic variant-effect predictors and recent deep learning approaches, including the pathogenic-variant predictor AlphaMissense and general protein language model–based methods. Finally, our results indicate that computational models can complement in vitro experiments in clinical variant interpretation, as StructureDCA predictions showed better agreement with clinically annotated phenotypes than large-scale deep mutational scanning data in several cases.

F. Pucci, Pauline Hermans, Matsvei Tsishyn et al. · 0 citations
Review Open access Aug 2026

Clinical Functional Assignment of TPMT and NUDT15 Alleles by the Clinical Pharmacogenetics Implementation Consortium Pharmacogene Curation Expert Panel

The Clinical Pharmacogenetics Implementation Consortium (CPIC) TPMT/NUDT15 Pharmacogene Curation Expert Panel (PCEP) conducted a comprehensive review of clinical, laboratory, and computational evidence to determine the clinical function assignments for TPMT and NUDT15 star alleles. These genes are critical for the metabolism of thiopurines, which are widely used in the treatment of cancer and autoimmune disorders. Standardized allele function assignment is essential for predicting metabolizer phenotypes and pharmacogenetics‐guided thiopurine dosing. The work presented here includes the first designation of decreased function alleles for both TPMT and NUDT15, reflecting new clinical data that demonstrate partial loss of enzymatic activity and reduced dose tolerance. The panel also reclassified several alleles previously assigned uncertain or unknown function. The functional assignments were informed by a standardized framework incorporating clinical data, such as thiopurine tolerance and toxicity, as well as in vitro protein activity, ex vivo enzymatic measurements, and in silico variant effect prediction tools. These updates enhance the precision of genotype‐to‐phenotype mapping and support more personalized thiopurine therapy across diverse patient populations.

Bailey M. Tibben, Maud Maillard, Victoria M. Pratt et al. · 0 citations
Open access Jul 2026

An integrated computational, clinical, and functional framework for assessing PTPN11 (SHP2) variant effects on ERK signaling and neural crest cell behavior in Noonan spectrum disorders

Germline mutations in PTPN11 cause Noonan syndrome (NS) and NS with multiple lentigines (NSML), yet how specific variants drive divergent clinical outcomes through distinct signaling and developmental mechanisms remains unclear. We find that germline and somatic mutations converge on N-SH2 and PTP domains but diverge at residue-level hotspots, reflecting distinct selective pressures. Clinical stratification of 18 pediatric patients reveals four distinct phenotypic classes including (i) the NSML-associated c.1403C>T (T468M) variant, characterized by lentigines, moderate growth impairment, and distinctive facial features; (ii) variants including the VUS c.1282G>A (V428M) and c.1432A>G (I478V), which were associated with cognitive deficits and variable growth impairment; (iii) c.1471C>A (P491T) and c.1472C>T (P491L), predominantly affecting cardiac and growth phenotypes with limited neurocognitive features; and (iv) a severe, multisystem class comprising c.172A>G (N58D), c.178G>A (G60S), c.844A>G (I282V), c.922A>G (N308D), and c.923A>G (N308S), spanning cardiac, growth, cognitive, and craniofacial abnormalities. Biochemical profiling in HEK293T cells revealed that PTPN11 variants stratify beyond simple gain/loss-of-function dichotomies into strong ERK-dependent hyperactivation, moderate ERK activation with variable protein stability and the paradoxical c.1282G>A variant, which did not increase ERK phosphorylation. In vivo, this variant drove excessive neural crest cell migration in chick embryos, suggesting that its effects on NCC migration may involve ERK-independent mechanisms or context-dependent signaling not captured by steady-state assays. ERK activation did not strictly correlate with clinical severity, yet these functional differences were associated with distinct growth, cardiac, pigmentation, and neurodevelopmental outcomes. Our data suggest lineage-specific sensitivity to SHP2 dosage, with dorsal root ganglia neurons appearing more vulnerable to reduced SHP2 stability than melanocyte precursors. Although direct correlations between specific signaling defects and individual clinical features remain complex, our findings provide a refined framework for PTPN11 variant classification, and reveal unexpected SHP2 functions in neural crest development.

M. Rodríguez-Martín, K. Cheriet, S. Adiba et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.