Skip to content

Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck.

Aug 2026 · Journal of Chemical Information and Modeling · Vol 66 16, pp. 9905-9915 · 0 citations · 50 references
Medicine

TL;DR

It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

Abstract

Virtual screening (VS) is an essential tool in drug discovery to prioritize potential drug candidates from vast chemical space. One key challenge limiting its performance is accounting for protein conformational flexibility. While ensemble docking methods have been developed to address this challenge by incorporating multiple protein conformations, these methods often rely on computationally intensive physics-based simulations to sample the relevant conformational space. Generative machine learning models offer a highly promising, scalable, and high-throughput alternative to overcome the limitations of these traditional approaches. We therefore investigate whether conformational ensembles generated by BioEmu, a recently developed generative model, can improve VS performance for kinase targets. Using the DUD-E benchmark data set and a validated AutoDock-GPU protocol, we generated and analyzed nearly 1300 structures across 26 kinases (approximately 50 structures each). BioEmu produces structurally diverse ensembles with substantial performance variation among individual structures. However, ensemble methods employing consensus or best-score selection fail to improve upon, and often degrade, VS performance compared to crystal structure baselines. To investigate the source of this limitation, we quantified the relationship between KinCoRe-based conformational state classification and screening performance. By calculating the coefficient of determination (R2) across the kinase subset, we found that the structural features governing VS performance differ substantially from those defining standard conformational states, with KinCoRe classifications leaving over 84% of performance variance unexplained. This critical gap demonstrates that structural diversity alone is insufficient to guarantee screening success. We show that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

View source

Similar papers

Open access Jul 2026

ScrambleBench: a workflow for comparative assessment of structure-based de novo generative models.

ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.

Veincent Yap, Pan Xu, Frankie S. Mak et al. · 0 citations
Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks.

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, Matthew B. Soellner, Charles L. Brooks · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations

MultiGeo: Predicting Drug-Target Affinity via Adaptive Multi-Conformation Ensemble Learning

MultiGeo is a DTA prediction framework that explicitly leverages multiple protein conformations rather than a single snapshot, and introduces a disagreement-aware gating mechanism that adaptively fuses this ensemble representation with the dominant structure only when the additional conformers provide complementary information.

Rui-Da Zeng, Cheng Guo, Yajie Meng et al. · 0 citations
Open access Jul 2026

Automated Parallel Synthesis Accelerates Virtual Screening Hit Discovery

Virtual screening (VS) is a powerful approach to exploring a vast chemical space, encompassing libraries of millions to billions of compounds. However, the low hit rates of VS require testing numerous candidates to validate true binders, followed by iterative optimization cycles, which makes experimental validation costly and time-consuming. Here, we report COMBINAUT, an automated parallel synthesis platform that generates diverse chemical scaffolds to accelerate hit validation and refinement. Using a faculty-wide collection of in-house building blocks, the system enables enumeration of over 22.9 million compounds, each designed for parallelized synthesis within 32 h using repurposed solid-phase peptide synthesis equipment. Using this platform, we performed large-scale VS targeting the allosteric pocket of the immuno-oncology target, C–C chemokine receptor 2 (CCR2). Our approach facilitated the rapid synthesis and testing of 100 VS hits spanning diverse molecular architectures. In radioligand binding assays, we successfully validated nine hits with distinct scaffolds, including completely novel CCR2 ligand chemotypes. Iterative hit-to-lead optimization using the automated workflow produced cell-active CCR2 antagonists. This work demonstrates the synergy of automated synthesis and VS, enabling the efficient exploration of chemical space and the rapid discovery of novel ligands.

Sean M McKenna, M. Šícho, Cas van der Horst et al. · 0 citations
Aug 2026

Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design

A state-aware functional classifier (SAFC) is developed that integrates molecular dynamics derived receptor ensembles, ensemble docking and protein ligand interaction graphs that provides dynamics-aware functional activity rankings for generated molecules that were partly complementary to docking, drug-likeness and synthetic accessibility scores.

H. Kumar, Zheng-Xiao Yang, Yankai Yu et al. · 0 citations