From benchmark accuracy to pharmacological credibility: why mechanistic explainability must define the next generation of AI for drug repurposing
Artificial intelligence (AI) is increasingly used in drug repurposing to integrate chemical, biological, and clinical data and to prioritize candidate drug–disease associations. Yet current evaluation practices remain dominated by benchmark metrics such as AUROC, AUPRC, F1, or top- k retrieval, which do not by themselves establish pharmacological credibility. A useful prediction is not merely one that scores highly on retrospective datasets, but one that can be connected to a biologically plausible mechanism, safety-relevant context, and a traceable evidentiary basis for experimental follow-up. Mechanistic explainability should therefore be treated as a core objective for translationally oriented AI-enabled repurposing, while also emphasizing that explainability complements rather than replaces experimental validation. Specifically, explanations should connect drugs, targets, pathways, phenotypes, and clinical outcomes in ways that are intelligible to pharmacologists and compatible with experimental prioritization, translational decision-making, and emerging regulatory expectations. We identify a central gap in the literature: many explainability methods emphasize feature attribution or local model transparency, but rarely produce pharmacology-aligned evidence chains that also communicate uncertainty, robustness, and provenance. We introduce biomedical knowledge graphs and neurosymbolic approaches as a promising foundation for more credible repurposing systems because they can support relational inference, structured mechanistic reasoning, and provenance-aware explanation. We further argue that future evaluation should extend beyond predictive accuracy to include mechanistic coherence, uncertainty, robustness, and provenance (MURP) for experimental pharmacology, ideally within workflows that integrate computational prediction with laboratory and, where feasible, real-world validation. Reframing success in this way would improve rigor, translational relevance, and regulatory readiness in AI-driven drug repurposing.