This work addresses one of the most pervasive obstacles to applying AI in real-world drug development by addressing conformal prediction framework tailored to label shift by weighting conformal scores using marginal label probability ratios and enhancing the trustworthiness of AI-driven predictions.
Abstract
Drug discovery and development underpins healthcare but remains costly and failure-prone. A critical bottleneck lies in predicting molecular properties such as solubility, potency, and toxicity, which directly determine whether a candidate can advance from preclinical to clinical trials. Artificial Intelligence (AI) has accelerated this process, yet its reliability is often undermined by distribution shift, as experimental conditions frequently diverge from training data. In addition, conventional point predictions provide only single-value estimates, offering limited guidance for high-stakes experimental design. We address these challenges with a conformal prediction framework tailored to label shift. By weighting conformal scores using marginal label probability ratios, our method produces statistically rigorous prediction intervals without retraining. This enables robust uncertainty quantification even when property distributions drift, directly tackling one of the most pervasive obstacles to applying AI in real-world drug development. By moving beyond accuracy alone to provide actionable confidence measures, our approach enhances the trustworthiness of AI-driven predictions. This further aligns predictive modeling with regulatory demands for transparency and uncertainty reporting and ultimately supports more reliable decision-making in billion-dollar development pipelines.
This scoping review investigates the current state of PK property prediction of small molecules in drug discovery using machine learning methods and a combination of machine learning and mechanistic models and proposes leveraging the pattern recognition capabilities of deep learning models in conjunction with the biological interpretability provided by mechanistic approaches.
Lucille Tomin, Vida Bodaghi-Namileh, D. Schwartz et al.· Journal of Chemical Informat...· 0 citations
This work presents a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail, and connects evaluation choices to real-world applications and case studies encountered in pharmaceutical research.
Srijit Seal, Akshat Shirish Zalte, David Alencar Araripe et al.· bioRxiv· 0 citations
A 4-layer framework can give MCOs earlier visibility into a therapy's likely clinical profile, cost-effectiveness distribution, and formulary placement probability before a manufacturer's dossier arrives, and allows budget forecasting and contracting strategy to keep pace with growing AI-accelerated pipelines.
Rishi Sharma· Journal of Managed Care & Sp...· 0 citations
Drug efficacy prediction remains a cornerstone of drug development and precision therapy. However, integrating heterogeneous biomedical data, including multi-omics profiles, pathological imaging, electronic health records, and pharmacokinetic-pharmacodynamic (PK/PD) time-series, faces three fundamental barriers, namely cross-domain distribution shifts between preclinical and clinical data, relational mismatches between isolated vector representations and biological networks, and feature heterogeneity across disparate modalities. To address these challenges, three AI paradigms have emerged, transfer learning for cross-domain alignment, graph neural networks for structured relational modeling, and Transformers for global cross-modal feature interaction. Importantly, these techniques form a many-to-many complementary system rather than a one-to-one correspondence, a key insight that this review explicitly formalizes. We further elaborate encoding workflows for PK/PD data to bridge static molecular signatures with dynamic in vivo exposure trajectories. Four graded clinical applications are outlined, including personalized monotherapy, combination optimization, drug repurposing, and preclinical-to-clinical evaluation of novel candidates. We also dissect persistent bottlenecks such as data harmonization, model interpretability, and prospective validation, and propose five actionable directions, namely privacy-preserving benchmarks, causally interpretable models, temporal dynamic frameworks, cross-domain generalization, and lightweight clinical tools. By integrating theoretical rationales, methodological synergies, and hierarchical translational scenarios, this review provides a unified roadmap to accelerate the clinical deployment of multimodal drug response prediction.
Jing-Wen Fang, Zihan Wang, Jun-Hao Shao et al.· Drug Discoveries & Therapeut...· 0 citations
Polypharmacy requires accurate prediction of drug-drug interactions to prevent adverse events, yet existing models often lack reliability and explainability. We propose T-DDI, a descriptor-based deep learning framework for multi-class drug-drug interaction prediction. Rather than relying on complex graph embeddings, T-DDI uses explicit physicochemical descriptors and an uncertainty-aware estimator to handle severe class imbalance. Evaluated on 868,069 drug pairs spanning 178 interaction types, T-DDI achieves a Macro F1 of 0.8452 on the held-out test set, improving to 0.8992 within the high-confidence subset (87.91% of test samples), outperforming all evaluated baselines within the architectures and datasets considered here. An illustrative prospective case-study assessment on five newly FDA-approved drugs from late 2025 showed that T-DDI can generate mechanistically plausible DDI hypotheses for compounds not used during model development. T-DDI pairs confidence-stratified predictions with LIME-based feature-level explanations and a web application for screening, supporting more reliable drug safety monitoring.
Q. Kha, Duc-Quang-Anh Nguyen, Phi Pham Van Hoang et al.· npj Digital Medicine· 0 citations
Evidence-constrained mechanistic synthesis (ECMS), a framework that classifies what information each finding contains and converts only that information into restrictions on a family of mechanistic hypotheses, is developed for the pre-calibration stage of drug development.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026