A transparent pre-screen can prioritise compounds ahead of structure-based calculation at a fraction of its cost, and the uncertainty and applicability-domain terms act as an abstention mechanism rather than an accuracy gain, and that abstention is not free.
Abstract
Background/Objectives: Molecular docking and molecular dynamics are accurate but computationally expensive, so the compounds entering them must be chosen well. The present study proposes CADT, a confidence-gated affinity–ADME-T docking-triage cascade that decides which compounds are worth docking. Methods: The gate combines an ensemble estimate of drug–target affinity with its epistemic uncertainty and an applicability-domain check. Predicted absorption, distribution, metabolism, excretion, and toxicity (ADME-T) developability is added as a soft flag. All components were trained on openly licensed Therapeutics Data Commons data. Ranking was assessed on the DAVIS and KIBA kinase panels and on BindingDB Kd, under three split protocols over five seeds. The routing decision was then examined against molecular docking, in which 407 compound–target pairs were docked into six withheld kinases. Results: A Morgan-fingerprint gradient-boosting model reached a concordance index of 0.866±0.006, with 0.813 for unseen targets and 0.720 for unseen drugs. Across eight ADME-T endpoints, the area under the ROC curve ranged from 0.65 to 0.91. On the cold-target split the cascade reduced the compounds sent to docking by 86% while retaining 61% of the true strong binders. Docking measured that reduction at 85%, and at an equal budget, the gate enriched true binders more than the docking score itself. Conclusions: A transparent pre-screen can prioritise compounds ahead of structure-based calculation at a fraction of its cost. However, the uncertainty and applicability-domain terms act as an abstention mechanism rather than an accuracy gain, and that abstention is not free.
A DeepPurpose model to estimate inhibition constants (Ki) from molecular graphs and protein sequences was trained and whether those predictions could be aligned empirically with measured IC50 values without treating Ki and IC50 as interchangeable was asked.
Basel Mansour, S. Dutta, Binil Benny et al.· bioRxiv· 0 citations
This study employed an integrated ligand-based virtual screening pipeline to identify potential CASP4 inhibitors from the DrugBank database, leveraging docking-score prioritization, SMILES-derived ChemBERTa embeddings, and key physicochemical descriptors. Building on this foundation, the workflow incorporated virtual...
Mubashir Hassan, S. Bhatti, Muhammad Yasir et al.· Journal of Chemical Informat...· 0 citations
DyAb is a pair-wise representation built on top of a pre-trained protein language model that achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across monoclonal antibodies targeting three different antigens.
J. Lin, Jennifer L. Hofmann, A. Leaver-Fay et al.· mAbs· 0 citations
Background and purpose: PIM-1 kinase, a serine/threonine kinase implicated in several cancers, has emerged as a promising yet underexplored target for anticancer therapy. This study aimed to identify the potential of PIM-1 inhibitors from marine natural products by integrating machine learning (ML)-based quantitative s...
Bishal Budha, Arjun Acharya, Madan Khanal et al.· Research in Pharmaceutical S...· 0 citations
Overall, the results show that classifier design and optimisation account for a substantial portion of the improvement over MolTrans, while the contribution of cross-modal architectural components is optimisation-sensitive and dataset-dependent.
Hao Pang, Fiseha B. Tesema, Tian-Xiang Cui et al.· BMC Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.