Biotoxins, mainly produced by venomous animals, plants, and microorganisms, exhibit high physiological activity and unique effects such as lowering blood pressure and analgesia. A number of venom-derived drugs are already available on the market, with many more candidates currently undergoing clinical and laboratory studies. However, drug design resources related to biotoxins are insufficient, particularly because of a lack of accurate and extensive activity data. To fulfill this demand, we developed the Biotoxins Database (BioTD). BioTD is the largest open-source database for toxins, offering open access to 14,607 data records (8,185 activity records), covering 8,975 toxins sourced from 5,220 references and patents across over 900 species. The activity data in BioTD are categorized into five groups: Activity, Safety, Kinetics, Hemolysis, and other physiological indicators. Moreover, BioTD provides data on 1,532 mutants, refines the whole sequence and signal peptide sequences of toxins, and annotates disulfide-bond information. All of the data in the database can be downloaded for free. Given the importance of biotoxins and their associated data, this new database is expected to attract broad interest from diverse research fields in drug discovery. BioTD is freely accessible at http://biotoxin.net/.
Gaoang Wang, Hang Wu, Yang Liao et al.· Journal of Chemical Informat...· 0 citations
Accurate prediction of blood-brain barrier permeability (BBBP) is essential for central nervous system drug discovery, yet existing models are often limited by their reliance on predefined physicochemical descriptors, small-molecule-centered training sets, or conformation-dependent representations, which restricts their transferability across chemically diverse modalities especially peptides. In addition, publicly available BBBP datasets remain fragmented, inconsistently standardized, and weakly controlled for molecular redundancy, increasing the risk of data leakage and overestimated model performance. In this study, we propose BBBP-Atlas, a structure-aware BBB permeability prediction model designed for unified modeling of small molecules and peptides with the first cross-modal dataset OmniBBBP. Designed to bypass descriptor and conformation dependencies, our model represents standardized molecular structures as atom-level graphs to capture local atom-bond environments and long-range topological dependencies associated with BBB transport. This design enables direct learning of structure-permeability relationships from molecular topology. For model training and evaluation, we curated a cross-modal, redundancy-filtered database OmniBBBP that seamlessly unifies small molecules and complex peptides, containing 10,218 unique compounds with 9,316 small molecules and 902 peptides. BBBP-Atlas achieved an accuracy of 0.8914 and an MCC of 0.7678 on the independent test set. On a balanced external benchmark of 200 compounds, our model reached an AUC of 0.9108, an accuracy of 0.8500, and an MCC of 0.7000, outperforming LightBBB by an absolute MCC gain of 6%. Case studies further showed that BBBP-Atlas captured clinically meaningful BBB permeability patterns, correctly identifying lorlatinib as BBB-permeable and vancomycin as BBB-impermeable with high confidence. The OmniBBBP-backed BBBP-Atlas offers a versatile and cross-modal approach for single-compound prediction, batch screening, and dataset exploration for CNS drug discovery. BBBP-Atlas is available at https://cadd.drugflow.com/bbbp/.
Xin Shen, Qun Su, Hao Luo et al.· bioRxiv· 0 citations
Targeting the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD) represents a promising strategy to overcome resistance in prostate cancer. However, its inherent lack of a stable tertiary structure and highly dynamic conformational ensemble pose formidable challenges for rational drug design. This study introduces an integrated computational workflow that combines enhanced sampling techniques and machine learning collective variables to identify druggable conformations of the AR-NTD and elucidate the binding mechanism of its modulator, EPI-002. We characterize nine metastable states of the Tau-5 region and reveal that ligand recognition is driven by π–π stacking and structured water-mediated hydrogen bonds. Leveraging these insights, we perform structure-based virtual screening based on the identified druggable conformations and identify K53, a rationally designed AR-NTD antagonist, which exhibits potent anti-proliferative activity in enzalutamide-resistant prostate cancer cells. K53 directly binds the AR-NTD, suppresses AR transcriptional activity, and demonstrates high selectivity for cancer cells. This work provides a rational design paradigm for targeting intrinsically disordered proteins and offers a therapeutic candidate for resistant prostate cancer. In this work, the authors develop a machine learning–based enhanced sampling workflow to target the intrinsically disordered AR-NTD, identifying druggable conformations and enabling transferable modeling of ligand binding for rational drug discovery.
Kai Zhu, Huating Wang, Jintu Zhang et al.· Nature Communications· 0 citations
ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in NSCLC and BRCA.
Yundian Zeng, Qing Ye, Jike Wang et al.· Journal of Chemical Informat...· 0 citations
The first comprehensive benchmarking framework specifically designed to accommodate inter-dataset heterogeneity is presented, finding that well-designed small datasets can match or even surpass the performance of larger benchmarks, suggesting that different metrics are applicable to different datasets/testing scenarios.
Yingjuan Cheng, Qing Ye, Linlong Jiang et al.· Journal of Cheminformatics· 0 citations
A novel committor learning framework grounded in the AlphaFold 3 paradigm is proposed that elucidates how ligand substituents regulate the ratio between distinct binding pathways, offering new perspectives for structure-based drug design.
Jintu Zhang, Zichang Jin, Huifeng Zhao et al.· 0 citations
The Comprehensive VS Platform with AI Engine (CVSP-AIE) for drug discovery from compound libraries integrates three AI models: KarmaDock, a fast docking model that directly updates atomic coordinates; CarsiDock, an accurate docking model that predicts protein-ligand distances and reconstructs binding poses; and RTMScore, an accurate scoring model that learns residue-atom distance distributions for affinity prediction.
Current structure-based drug design generative models often struggle to faithfully recapitulate genuine ligand-protein binding interactions. Instead, under the coupling of implicit learning architectures and biased training data, they tend to learn spurious statistical correlations. To address this, we propose EIP-Diff (Explicit Interaction-Prompted Diffusion), an architecture featuring a novel explicit interaction-prompt embedding mechanism that is better suited for real-world target-specific drug design. This architecture replaces biased implicit learning with explicit, residue-level biological guidance, thereby promoting more fine-grained geometric fidelity and more precise interaction-aware conditioning. To fully realize the capabilities of EIP-Diff and provide a reliable basis for performance evaluation, we further constructed CrystalData set, which provides higher-fidelity and less-biased structural supervision than existing data sets. This explicit architecture markedly improves distribution consistency: even when trained on the crossdocked data set, EIP-Diff achieves the highest alignment with authentic pharmacological distributions among evaluated models. Training on CrystalData set further enhances this alignment and improves 3D geometric accuracy, while retaining strong controllability, high chemical space coverage, and near-perfect uniqueness. In addition, target-based validation on KAT6A and YTHDC1 confirmed that EIP-Diff accurately recapitulates native-like binding modes. Furthermore, in a real-world drug design task against IDO1, we successfully designed a novel lead compound with nanomolar potency (IC50 = 0.31 nM). These results demonstrate that the EIP-Diff architecture can explicitly leverage experimentally derived structural data and biologically meaningful interaction information for target-specific molecular generation, thereby enabling its effective application to real-world structure-based drug design.
Huabin Du, Mingyang Wang, M. Luo et al.· Journal of the American Chem...· 0 citations
This paper introduces Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language, and incorporates a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering.
Mingze Yin, Yiheng Zhu, Jialu Wu et al.· Proceedings of the 32nd ACM...· 0 citations