Skip to content
Open access

Improving Generalizability in Whole-Cell Antibiotic Discovery Through Active Learning

Jul 2026 · bioRxiv · 0 citations · 66 references
Medicine Biology

TL;DR

These results demonstrate that calibrated AL strategies can overcome data acquisition bottlenecks and train generalizable property predictors able to extrapolate to OOD molecules.

Abstract

Machine learning (ML) has accelerated molecular discovery, yet training models to generalize to out-of-distribution (OOD) chemical spaces remains fundamentally constrained by the high cost of experimental validation. In antibiotic discovery, where whole-cell phenotypic high throughput screening (HTS) is resource-intensive, iterative ML-guided compound selection – or Active Learning (AL) – offers a pathway to efficiently navigate available chemical spaces. However, the algorithmic tradeoffs between prioritizing compound novelty (exploration), predicted bioactivity (exploitation), and their impact on OOD generalizability remain unresolved for noisy, whole-cell biological systems. In this work, we systematically evaluate three AL strategies for whole-cell bacterial bioactivity and benchmark their effects on model accuracy, hit rate, and OOD performance. Using retrospective simulations on Mycobacterium tuberculosis HTS data, we identify an optimal AL strategy that balances predicted hit/non-hit novelty with overall hit rate. We then integrate the strategy in a closed-loop Borrelia burgdorferi antibiotic discovery HTS campaign. The AL-guided approach successfully increased the experimental screening hit rate five-fold (from a 0.2% rate within investigator-selected plates to 1.0%). Further, when the trained model was applied in prospective in silico selection of highly diverse compounds across multiple bacterial species, the AL-trained whole-cell inhibition predictor demonstrates 53-fold enrichment over investigator-directed screening (11.0% experimental validation of predicted hits). Of these, 100% demonstrated the intended narrow spectrum activity for Borrelia burgdorferi. These results demonstrate that calibrated AL strategies can overcome data acquisition bottlenecks and train generalizable property predictors able to extrapolate to OOD molecules.

Read PDF

Similar papers

Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks.

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, Matthew B. Soellner, Charles L. Brooks · 0 citations
Open access Jul 2026

Smiles-based bioactivity prediction through molecular encoder selection and data augmentation.

Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for "Jump AI(.py) 2025: 3rd AI Drug Discovery Competition", hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder-regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.

Ju Hyung Lee, S. Choi, Utku Ozbulak et al. · 0 citations
Review Aug 2026

Machine learning for precision prediction of antimicrobial peptide activity and spectrum.

The accelerating global crisis of antibiotic resistance demands new therapeutic paradigms, and antimicrobial peptides (AMPs) have emerged as promising candidates owing to their broad activity and reduced propensity for resistance development. However, despite rapid progress in AMP discovery and generation, the accurate prediction of antimicrobial potency and activity spectrum remains a major bottleneck for clinical translation. In this Review, we examine how recent advances in machine learning are reshaping AMP research, driving a shift from large-scale discovery toward precision-guided prediction and design. We first summarize the molecular mechanisms underlying AMP function and critically assess existing AMP databases from the perspective of machine learning readiness, highlighting limitations in quantitative and spectrum-resolved annotations. We then review recent developments in peptide representation learning, describing how modern models encode sequence, structure, and dynamic features to capture antimicrobial activity. Building on this foundation, we discuss progress in de novo AMP design and emerging frameworks for quantitative minimum inhibitory concentration prediction and strain-specific spectrum profiling. Finally, we outline future directions for the field, emphasizing integrated generative-predictive pipelines, interpretable models, and closed-loop experimental validation as key enablers for the development of potent, selective, and clinically viable antimicrobial therapeutics.

Tianxiao Wan, Yiling Wang, Tianle Ren et al. · 0 citations
Review Aug 2026

Advancing cancer drug discovery through the integration of machine learning and high-throughput screening.

This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization.

K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al. · 0 citations
Review Jul 2026

In Silico ADMET: From Current Practices to Novel Profilers.

Multitask learning is a promising strategy in computational drug discovery, potentially improving predictive performance and generalization over traditional single-task models. MTL has shown particular value in absorption, distribution, metabolism, elimination, and toxicity (ADMET) and potency predictions, which are key for drug design. Yet, many existing Web servers rely on the same uncurated, decade-old data sets, creating an illusion of diversity. This work critically reviews open-source ADMET Web services, revealing extensive data redundancy and limited curation across the field. We introduce OneADMET, a meticulously curated data set of 738,161 compounds with 1,119,719 measurements spanning 44 ADMET end points and 1 489 biological activities. We report a unified ChemProp-based MTL model capable of handling hundreds of continuous tasks simultaneously, which has practical advantages for model deployment and maintenance. Additionally, we observed that these MTL models match or surpass single-task models in predictive accuracy. This study highlights the utility of large-scale MTL for pharmacokinetics profiling and contributes practical tools and data sets for the community.

P. Llompart, C. Minoletti, G. Marcou et al. · 0 citations
Open access Jul 2026

Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

A machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round that may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants.

Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al. · 0 citations