PandaDock’s empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR.
Abstract
We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docking, and an SE(3)-equivariant graph neural network scoring function trained at scale. Ligand flexibility is represented as a torsion tree and pose parameters are optimized by Monte Carlo with Metropolis acceptance refined by L-BFGS, with rotational gradients obtained in closed form through the derivative of the SO(3) exponential map rather than by finite differences. Affinity grids are built by a blocked neighbor-selection scheme that is exact and 5.6–9.7× faster than dense evaluation, and may be cached across ligands sharing a receptor and site, reducing a six-ligand series from 29.3 s to 10.4 s. On 814 protein–ligand complexes spanning 14 target families, PandaDock recovers a pose within 2 Å of the crystal geometry in 33.7% of cases at rank 1 and in 57.0% of cases within the returned ensemble. The GNN scoring function is trained on 741,706 co-folded complexes from SAIR under target-disjoint splits, reaching a Pearson r of 0.407 on 90,219 held-out complexes and transferring to 202 independent crystal structures with measured Ki, Kd, IC50 or EC50 at r = 0.467. We report the model against three controls, a target-mean predictor, a ligand-descriptor-only baseline, and within-target correlations, and document both where it performs and where it does not, including its unsuitability for pose rescoring. On an independent 30-compound series against a single GABAA receptor target, PandaDock’s empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR. At full scale on the PDBbind v2020 refined set (n = 4,640, native crystal poses), the fully independent SAIR model reaches r = 0.531, and a dedicated model trained on PDBbind alone under a target-disjoint split reaches r = 0.690 on its own held-out test complexes – the strongest evidence in this work that PandaDock’s affinity predictions generalize. PandaDock is distributed under an open-source license at https://github.com/pritampanda15/PandaDock with a complete command-line interface and a reproducible benchmarking harness.
This work introduces Evolutionary Driven Bayesian Optimization (EA-BO), a surrogate-based framework designed for efficient exploration under strict evaluation budgets and demonstrates that EA-BO provides data-efficient strategy for locating promising docking regions when computational cost limit traditional approaches.
A. Lopez-Rincon, B. Varga, D. Rojas-Velazquez et al.· Annual Conference on Genetic...· 0 citations
Physics-based protein–ligand docking critically depends on efficient pose sampling, yet established sampling and local refinement algorithms can be inefficient and unstable in the highly nonconvex energy landscapes characteristic of protein–ligand interactions. To address this limitation, we introduce an enhanced local optimization strategy based on curved line search (CLS) and integrate it into AutoDock Vina, resulting in Vina_CLS. The proposed method enables more flexible step-size selection during local refinement and improves convergence in challenging regions of the energy landscape. Across benchmarks on the PDBbind refined set and the LEADS-PEP data set, Vina_CLS consistently outperforms the baseline, exhibiting greater robustness by solving more docking problems, as well as improved efficiency through reduced function and gradient evaluations and shorter runtimes. These gains translate into practical benefits, including more frequent identification of difficult-to-access local minima, enhanced redocking accuracy, and increased recovery of near-native poses. Together, these results demonstrate that improved local optimization can substantially enhance docking performance, highlighting an important, underexplored opportunity to advance structure-based drug discovery.
Leo Gaskin, Matthias Welsch, J. Kirchmair et al.· Journal of Chemical Theory a...· 0 citations
A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.
Ryan Varghese, Pooja Tiwary, Krishil Oswal· bioRxiv· 0 citations
The kinetics of protein-ligand binding systems are increasingly recognized as a key determinant of drug efficacy, yet remain far harder to compute than binding affinities. Existing kinetics methods either bias the dynamics along a collective variable (CV), demanding careful system-specific CV design, or use path sampling, which keeps the dynamics unbiased but can struggle to converge rates out of deep free-energy wells and often relies on hand-engineered descriptors. By combining the `best of both worlds', we propose a method to compute accurate kinetics for general ligand-unbinding problems at modest computational expense and minimal fine tuning, building on the AI for Molecular Mechanism Discovery (AIMMD) path sampling framework. To avoid the need for feature engineering, we opt for modelling the committor with a single descriptor-free, equivariant graph neural network shared across all systems. We also partially flatten deep bound-state wells with a static, basin-restricted bias potential. This improves convergence by lifting the path sampling state boundary out of regions, where the committor is hard to learn, while leaving the reactive region strictly unbiased. Across host-guest and protein-ligand systems spanning roughly 17 orders of magnitude in residence time, the method robustly recovers rates in line with reference and experimental values. Simultaneously, and without further sampling, it also reconstructs the underlying unbinding mechanisms. We additionally find that accurate rates do not require globally accurate committor models, allowing for efficient kinetics estimation even in a low-data training regime. Requiring little system-specific setup, our approach offers an efficient and broadly generalizable route to binding kinetics, and its shared committor architecture lays crucial groundwork for probing structure-kinetics relationships across ligand series in drug discovery.
In 2025, we released UAM-Ixachi to democratize, simplify, and accelerate molecular docking and virtual screening methods. It is a free, open-source, and user-friendly tool. Building on that work, we present SMASH, which features several upgrades: it predicts binding sites using machine learning, automatically determines titration states, utilizes a Graphics Process Unit for molecular docking, combines machine learning with energy scoring functions, and clusters data using the K-means algorithm. Tests demonstrated that the tool might handle large projects within a reasonable time using local computational resources. In automatic mode, the tool can accurately reproduce a high percentage of ligand poses from Protein Data Bank complexes via docking simulation. It also differentiated between the predictions of active and decoy ligands from a DUD-E data set in a reasonable amount of time. SMASH is freely available at https://smashreleases.z13.web.core.windows.net/ We implemented several software solutions to prepare and execute molecular docking simulations: PDB2PQR, P2Rank, MGL Tools, OpenBabel, AutoDock-GPU, Vina-GPU, and SCORCH. The MMFF94, AD4, and Vina force fields are part of the implemented tools.
A. Suárez-Alonso, L. D. Herrera-Zúñiga, Mayra Lozano-Espinosa et al.· Journal of Molecular Modelin...· 0 citations
Gaussian accelerated molecular dynamics (GaMD) enhances conformational sampling by adding a smooth boost potential without requiring predefined collective variables, but an engine-integrated implementation has not been available in GROMACS. Here, we implement total-, dihedral-, and dual-boost GaMD in GROMACS 2025.4, including staged energy-statistics collection, GPU-based bias evaluation and force scaling, restart support, and outputs required for cumulant-based free-energy reweighting. The implementation was evaluated using four benchmark systems spanning conformational free energies, protein folding, and ligand recognition. For alanine dipeptide, a reweighted 100 ns GaMD trajectory recovered the major free-energy basins and rotational barriers in overall agreement with a 1000 ns conventional MD simulation. For chignolin and TC5b, all three independent trajectories for each system sampled native-like folded states from extended conformations within 300 ns and 1 μs, respectively; the best TC5b structure had a minimum backbone RMSD of 0.03 nm from the experimental structure. In the benzene–T4 lysozyme system, two of five independent 500 ns trajectories captured both ligand binding and dissociation, yielding a bound pose with a minimum ligand RMSD of 0.06 nm from the crystal structure. Across all four systems, the boost-potential distributions were approximately Gaussian, and second-order cumulant reweighting resolved the expected conformational and binding free-energy basins. These results demonstrate that GROMACS-GaMD provides a practical, GPU-enabled, collective-variable-free enhanced-sampling framework for biomolecular free-energy calculations, protein folding, and ligand-binding studies.
Yuefeng Yang· bioRxiv· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.