Skip to content
Open access

ConfDock: Atom-specific Uncertainty Quantification for Molecular Docking via Conformal Prediction

Aug 2026 · bioRxiv · 0 citations · 72 references
Biology

TL;DR

Results demonstrate that combining learned, structure-aware quantile estimation with conformal calibration enables rigorous uncertainty quantification for molecular docking at atom-level resolution.

Abstract

Molecular docking is widely used in structure-based drug discovery, yet most approaches provide point estimates without rigorous uncertainty quantification. This limitation makes it difficult to assess when a predicted pose should be trusted, especially when docking methods are applied to diverse protein–ligand systems. We present ConfDock, a conformal prediction (CP) framework for constructing atom-specific prediction intervals for ligand docking poses. ConfDock combines graph neural network (GNN) based quantile estimation with split conformal calibration, producing intervals that adapt to local protein–ligand environments while retaining distribution-free finite-sample coverage guarantees. We evaluate ConfDock on 238 protein–ligand complexes across four docking methods representing distinct computational paradigms. The proposed approach yields substantially narrower prediction intervals compared to standard split CP (57.2% average reduction in mean interval width, up to 74.5%) while maintaining target coverage across all evaluated settings. Ablation analysis indicates that the GNN captures the dominant structure-dependent variability in uncertainty, whereas the conformal calibration step provides a bounded adjustment to ensure coverage guarantees. These results demonstrate that combining learned, structure-aware quantile estimation with conformal calibration enables rigorous uncertainty quantification for molecular docking at atom-level resolution.

Read PDF

Similar papers

Book Jul 2026

Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion

Accurate protein–ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust. Consensus scoring and ensemble methods improve mean accuracy but treat all predictions identically without interpretable confidence measures or uncertainty decomposition, ignoring the chemical context of each protein–ligand pair. To address this limitation, we introduce RELIABLE-BA (RELIABiLity-aware Evidential fusion for Binding Affinity), an evidential framework for multi-engine binding affinity prediction. Our model comprises three steps: (1) modeling each engine as an evidential expert via Normal–Inverse-Gamma distributions, (2) scaling epistemic uncertainty through learned reliability from molecular context while preserving each expert's predictive mean, and (3) fusing experts through closed-form aggregation that captures both individual uncertainty and inter-engine disagreement. Experiments on the PDBBind and BDB2020+ benchmarks demonstrate competitive point prediction with substantially improved uncertainty calibration, and additional validation on the SARS-CoV-2 Mpro dataset and 5HT2A receptor demonstrates applicability to clinically relevant drug targets. Crucially, these uncertainty estimates enable reliable filtering of protein-ligand pairs, reducing prediction error by up to ~25% when retaining only high-confidence pairs. To our knowledge, RELIABLE-BA is the first multi-engine binding affinity prediction framework to combine evidential fusion with context-dependent reliability, offering a principled path toward trustworthy AI-guided drug discovery. Our code is publicly available at https://github.com/yongchand/RELIABLE-BA.

Yongchan Hong, Defu Cao, Wenjin Liu et al. · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations
Jul 2026

Atomic Uncertainty Pinpoints Critical Failure Structures for Trustworthy Molecular Property Prediction.

Accurate prediction of molecular properties is essential for accelerating drug discovery, yet current deep learning methods generally lack reliable uncertainty estimates, particularly at the atomic level. A limitation of current approaches is the reliance on global molecular uncertainty, which frequently masks localized ambiguities by averaging the uncertainty across the entire molecule. To address this, we introduce AUINet, an atomic uncertainty-aware iterative network that quantifies and refines uncertainty at the atomic scale. Built on a D-MPNN architecture, AUINet uses Monte Carlo dropout to estimate atom-level uncertainty and iteratively refines atomic features through uncertainty-guided updates. Comprehensive evaluations show that AUINet outperforms state-of-the-art models on molecular property benchmarks and protein-protein interaction inhibitor tasks under low-data conditions, all without requiring extensive pretraining. Crucially, rejection sampling experiments reveal that atom uncertainty provides a more robust error signal than classic molecular uncertainty. More importantly, AUINet provides chemically interpretable insights by pinpointing specific functional groups and structural motifs that contribute most to prediction uncertainty, as validated in solubility prediction and activity-cliff analysis. Overall, AUINet's precise localization of atomic-level uncertainty establishes a new paradigm for trustworthy molecular property prediction.

Jiayu Qian, Yukun Luo, Qingping Zhou et al. · 0 citations
#machine learning Preprint Aug 2026

Conformal Prediction for Molecular Properties under Label Shift

This work addresses one of the most pervasive obstacles to applying AI in real-world drug development by addressing conformal prediction framework tailored to label shift by weighting conformal scores using marginal label probability ratios and enhancing the trustworthiness of AI-driven predictions.

Hyeonsu Lee, Juyeong Kim, Erkhembayar Jadamba et al. · 0 citations
Book Open access Jul 2026

Evolutionary‑Driven Bayesian Optimization for Automated Molecular Docking with AutoDock Vina

This work introduces Evolutionary Driven Bayesian Optimization (EA-BO), a surrogate-based framework designed for efficient exploration under strict evaluation budgets and demonstrates that EA-BO provides data-efficient strategy for locating promising docking regions when computational cost limit traditional approaches.

A. Lopez-Rincon, B. Varga, D. Rojas-Velazquez et al. · 0 citations