Boltz2ESI is introduced, an end-to-end framework that predicts enzyme–substrate interactions by leveraging structural knowledge learned by a biomolecular foundation model and consistently outperforms state-of-the-art sequence-based and rigid-docking approaches.
Abstract
Enzymatic catalysis relies on precise structural and chemical complementarity, yet systematically mapping enzyme-substrate interactions remains a critical bottleneck. While structure-aware methods have advanced functional annotation, their reliance on predefined binding pockets and rigid-body docking fails to capture the ligand-induced conformational changes essential for catalytic turnover. Here we introduce Boltz2ESI, an end-to-end framework that predicts enzyme–substrate interactions by leveraging structural knowledge learned by a biomolecular foundation model. Through native co-folding, the framework inherently captures active-site plasticity without requiring predefined pocket annotations. Integrating these learned biophysical priors with global evolutionary context and geometric molecular descriptors, Boltz2ESI consistently outperforms state-of-the-art sequence-based and rigid-docking approaches. Extensive validation demonstrates that the framework accurately discriminates tight sub-family specificities, enabling effective candidate prioritization for biosynthetic pathway elucidation, as demonstrated on the withanolide pathway. Ultimately, this structure-dynamic approach establishes an actionable foundation for accelerating rational biocatalyst discovery and large-scale pathway de-orphaning.
Whether binding specificity and partner selection in protein-protein interactions (PPIs) can be reliably inferred from static structures or require more dynamic, pathway-resolved energetic analyses remains an open question. To explore this, we focus on the ornithine decarboxylase (ODC)-antizyme isoform 1 (Az1)-antizyme inhibitor (AzIN) system, a well-characterized competitive PPI network that plays a critical role in regulating polyamine homeostasis. By combining extensive all-atom molecular dynamics simulations with biochemical experiments and the development of a new tool, we uncover key dynamic features of the static and recognition pathway interaction. Based on these, we designed novel antizyme isoforms (NAZs). Our analysis, using residue-resolved energetic landscapes, reveals critical determinants of binding specificity and partner selection that static structures alone cannot capture. These insights guide the engineering of NAZs that either directly engage ODC or modulate Az1 availability. This work provides a new perspective, demonstrating that dynamic energetic landscapes, rather than static structures, are key to understanding and modulating competitive protein recognition. Additionally, our DyResEL tool enables broader, more detailed analyses of energetic contributions, offering a versatile approach for exploring PPIs in various biological contexts.
Baolin Guo, Qian Xue, Fan Yang et al.· Journal of Chemical Informat...· 0 citations
Current rational drug design relies predominantly on computational (CADD/AIDD) methods that model binding thermodynamics and static conformations of target proteins, primarily in their inactive states. However, the kinetic parameters that govern experimental efficacy—such as catalytic turnover and signaling potency—are determined by molecular interactions with transition states (TS), intermediate states (IS), and the entire continuum of conformations along the least free-energy activation pathway. The absence of this dynamic dimension has fundamentally limited the predictive power and success rate of conventional structure-based approaches. Here, we present a structural database that systematically maps the complete activation trajectories of pharmaceutically relevant targets, encompassing TS, IS, and all connecting conformational ensembles. This resource offers multiple strategic advantages for drug discovery: enabling rational targeting of previously “undruggable” proteins, facilitating biased agonism/antagonism design, revealing cryptic allosteric sites in inactive conformations, identifying novel transient pockets along the activation route, rationalizing the mechanisms of existing drugs, predicting mutational effects on activation barriers, and prospectively forecasting drug resistance and off-target liabilities. We demonstrate the utility of this database through representative case studies and provide implementation guidelines for integration into existing discovery pipelines. More detailed information can be found at our website: https://www.momedpamdb.com/en. Terminology The following terms are clarified in this document: Stable state (SS): In this document, this term refers exclusively to, and is synonymous with, the protein’s inactive state (IAS). Note that other states may also be stabilized into meta-stable states by certain means. Unstable state (US): This term encompasses all states other than SS, even if they appear computationally meta-stable on the free energy surface. Activated state (AS): The meta-stable working state of the protein. Transition state (TS): The state with the highest free energy along the least-energy pathway on the free energy surface that connects the inactive state to the activated state of the target protein. Intermediate state (IS): The state(s) located at a local minimum along the least-energy pathway, excluding SS and AS.
Co-folding models hold immense potential for allosteric drug discovery, but have been severely hampered by their systematic bias toward orthosteric ligand binding. While fragment screening has been proposed for allosteric binding site discovery, we show that co-folding models still suffer from memorization in which chemically simpler fragments also default to canonical orthosteric binding sites. To overcome these limitations, we introduce CAFE (Co-folding Approach for Fragment Exploration), a co-folding protocol that uses competitive orthosteric blockers to divert fragments into non-canonical sites as illustrated here with the Boltz-2 co-folding model. Using ADP as an orthosteric blocker for the kinase family, we find CAFE substantially increases the allosteric binding site exploration for fragments, with notably strong absolute binding free energies that match or exceed those of known crystallographic poses, without post-hoc refinement of the Boltz-2 prediction. We also show that CAFE identifies cryptic binding pockets undetected by conventional pocket prediction tools, some of which are more thermodynamically favorable than the allosteric or orthosteric pockets. To demonstrate generality, we apply CAFE using Type I orthosteric blockers for kinase proteins, known orthosteric ligands as blockers for non-kinase proteins in the RAS-MAPK signaling pathway, and for virtual screening campaigns using fragment libraries for new fragments that selectively engage allosteric and cryptic binding sites. CAFE establishes orthosteric blocking and fragment screening as a training-free, inference-time protocol that helps overcome some of the limitations of current co-folding models while elevating their great promise for allosteric and cryptic binding drug discovery.
Justin Purnomo, Kunyang Sun, T. Head-Gordon· bioRxiv· 0 citations
Understanding the mechanisms underlying large scale protein conformational changes in signaling pathways is critical for elucidating disease processes and developing targeted therapeutics. However, existing experimental and computational methods struggle to resolve the dynamic ensembles of intermediate states that mediate such transitions, particularly in large biomolecular complexes. Here, we introduce a two-stage generative diffusion modeling framework designed to support pathway discovery in protein complexes, demonstrated using RAF kinase dimerization, a key event for kinase activation and oncogenic signaling. Our approach first generates ultra-coarse-grained structures conditioned on low dimensional descriptors along the monomer-to-dimer transition. It then applies a super-resolution model to recover detailed coarse-grained topologies suitable for molecular simulation. We show that this framework produces physically plausible, diverse, and robust intermediate structures, even for previously unseen interpolated descriptor values. The resulting ensemble enables generation of closely spaced candidate intermediate structures between biophysically distinct states, providing valuable starting points for downstream adaptive sampling and mechanistic studies. Overall, our results highlight the potential of diffusion-based generative models to bridge the gap between static structural data and isolated ensembles, and the dynamic complexity of protein signaling pathways.
Tim Hsu, Konstantia Georgouli, Michael Jones et al.· Machine Learning: Science an...· 0 citations
Protein phosphatase 2 A (PP2A) achieves signaling specificity through regulatory B subunits, but the chemical and structural determinants of regulatory-subunit recognition surfaces remain incompletely defined. The first PP2A-B55α complex structure identified a regulatory groove on the β-propeller surface, spatially separated from the catalytic site and occupied by a FAM122A regulatory segment. This groove therefore represents a macromolecular recognition surface that can be systematically probed for cyclic peptide engagement. Here, 8466 cyclic peptides from CycPeptMPDB were screened against the B55α regulatory groove, and KarmaDock score-based ranking prioritized seven representative cyclic peptides for detailed structural and energetic analysis. Refined simulations showed peptide-dependent modulation of conformational stability, convergence to stable bound states, selective stabilization of the regulatory groove, and retained flexibility of the extended A-subunit arm. Triplicate and extended simulations of P-659, together with triplicate simulations of the top candidate P-589, further supported reproducible structural behavior and binding energetics. Persistent hydrogen-bonding patterns suggested peptide-specific anchoring through Asp190, Asp197, Asp333, Tyr330, Ser280, and Lys345. Peptide binding was accompanied by the expected displacement of solvent from the solvent-exposed groove interior and localized reorganization of interfacial hydration, while residue-wise thermodynamic profiling identified Lys81, Met215, Glu216, Phe273, Tyr330, Asp333, and Phe336 as key solvent-response residues. Alanine scanning identified Asp197 as the principal energetic hotspot, with ligand-specific contributions from Ser280, Tyr330, and Asp333. Binding free-energy calculations indicated balanced gas-phase and solvation contributions, with P-589 showing the most favorable relative MM-GBSA binding-energy estimate among the analyzed peptides (ΔGTOTAL = -61.84 ± 0.43 kcal/mol). Together, these data establish a computational, structure-, dynamics-, and energetics-resolved framework for cyclic peptide recognition at the PP2A-B55α regulatory groove and define testable hypotheses for experimental validation.
Muhammad Waqas, Li Xuan, Haoke Zhang et al.· International Journal of Bio...· 0 citations