Skip to content
Preprint

HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration

Jul 2026 · 0 citations · 86 references
Computer Science

TL;DR

HEDGEHOG is introduced, a unified six-stage filtration benchmark that is inspired by industrial hit identification workflows and exposes a central limitation of current molecular generators: molecules that appear acceptable under isolated criteria rarely satisfy medicinal chemistry, synthesis, docking, and 3D pose filters simultaneously.

Abstract

Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance target-relevant activity, synthetic accessibility, physicochemical properties, and other multiparameter design constraints. However, metrics commonly used to evaluate molecular generators only weakly reflect whether the generated compounds are medicinally plausible and suitable for downstream computation. This can produce false positives in model evaluation, incorrect assumptions, and inefficient use of computational resources. We introduce HEDGEHOG, a unified six-stage filtration benchmark that is inspired by industrial hit identification workflows: (i) preprocessing; (ii) physicochemical descriptor screening; (iii) structural alerts and graph-sanity checks; (iv) synthesis feasibility; (v) docking and binding affinity estimation; and (vi) three-dimensional pose and interaction checks. We evaluate 23 molecular generators across three model classes under a standardized protocol. Across 230,000 generated molecules, only 0.65% of initial molecules survive all stages. Our results expose a central limitation of current molecular generators: molecules that appear acceptable under isolated criteria rarely satisfy medicinal chemistry, synthesis, docking, and 3D pose filters simultaneously.

View source

Similar papers

Open access Jul 2026

ScrambleBench: a workflow for comparative assessment of structure-based de novo generative models.

ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.

Veincent Yap, Pan Xu, Frankie S. Mak et al. · 0 citations
Preprint Aug 2026

MolecularCanvas: LLM-assisted Small-Molecule Drug Discovery via Structure-Guided Constraints

MolecularCanvas is an interactive system that enables users to iteratively construct an optimization context by integrating high-level goals, structure-level annotations, property constraints, and reference-based preferences that guides the generation of candidate molecules across diverse molecular structures.

Haoyu Dong, Rui Sheng, Shuhao Zhang et al. · 0 citations
Aug 2026

Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design

A state-aware functional classifier (SAFC) is developed that integrates molecular dynamics derived receptor ensembles, ensemble docking and protein ligand interaction graphs that provides dynamics-aware functional activity rankings for generated molecules that were partly complementary to docking, drug-likeness and synthetic accessibility scores.

H. Kumar, Zheng-Xiao Yang, Yankai Yu et al. · 0 citations
Aug 2026

Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck.

It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

Jaeoh Shin, K. Joo, Jejoong Yoo · 0 citations
Open access Jul 2026

Strategies for Identifying Molecules of Interest in Large Chemical Spaces

In recent years make-on-demand compound libraries (so-called Chemical Spaces) have gained more and more interest in pharmaceutical industry. Compound vendors promise cheap compounds, fast delivery, high synthetical accessibility and a large pool of novel chemical matter, fulfilling the requirements of fast Design-Make-Test cycles. Searching in ultralarge Chemical Spaces with known 2D similarity metrics, like fingerprint-based Tanimoto, substructure, or pharmacophore similarity searches, contains pitfalls due to the representation of molecules as synthons with connectivity rules. Applied to a set of almost 3000 drug-relevant queries we analyzed the ability of similarity search methods to retrieve analog compounds from Chemical Spaces, and how to best approach typical use cases in early phase drug discovery. Distinct characteristics of each similarity metric suggest orthogonal complementarity, enabling a versatile framework to diverse challenges present in hit discovery and lead expansion campaigns. Our investigations resulted in formulating practical considerations and guidelines for interpreting the scores including recommendations for thresholds for each similarity metric.

Raphael Klein, Sascha Jung, Alexander Neumann et al. · 2 citations
Open access Jul 2026

Prioritising search for virtual screening via preliminary interpretable low-feature likelihood-based rankings of drug-target activity measures.

BACKGROUND Current AI-based Virtual Screening (VS) methods seek to manage ultra-large molecular libraries. To this end, they develop increasingly efficient heuristics to rank ligands by their predicted activity against a target protein. However, these methods remain computationally demanding due to the billion-scale compound libraries that must be evaluated without prior, informed guidance. RESULTS This article proposes an offline/online method that: (1) Wisely selects (once and forall, offline phase) a small number of easy to compute features [Formula: see text] of both the amino acid sequence of the proteins ([Formula: see text]) and the molecular structure of the ligands ([Formula: see text]), and discretises their domains; this induces a low-dimensional finitisation of proteins' and ligands' chemical spaces. (2) Given a target protein [Formula: see text], immediately returns (online phase) a likelihood-based ranking of the classes of the ligands' chemical space, in descending order of the estimated probability that molecules in each class will achieve a satisfactory activity measurement against [Formula: see text]. This enables any VS method to prioritise the search to the most promising subsets of candidates. To ensure statistically robustness, our offline feature selection: (a) leverages knowledge stemming from a huge dataset of 2 559 403 entries (ligand-protein activity measurements) obtained by unifying the most representative sources regarding biochemical kinetics (Brenda, Sabio-rk, BindingDB) and augmented with 3781 features computed by 7 well-known third-party software tools; (b) explicitly aims at low-dimensional coarse-domain feature spaces; (c) takes proper countermeasures to prevent biases in the source data and overfitting; (d) supports iterative improvement of [Formula: see text] via an anytime offline algorithm and means to interactively exclude features deemed uninformative upon rankings inspection; (e) supports intepretability of the rankings by enabling inspection of the features' values characterising each ligand class. CONCLUSIONS By evaluating our rankings on evaluation data (from PDBbind and additional BindingDB entries unsuitable for accurate statistical analysis), we demonstrate their effectiveness for library prioritisation. Specifically, our findings indicate that approximately 60% of the high-affinity ligands occur in the top 25% ranked ligands' classes, while 85% fall within the top 50%. Furthermore, we conduct retrospective analysises using AutoDock Vina scores for over 260 000 molecules across 58 medically relevant targets. Results demonstrate that our method cuts the number of dockings needed to retrieve an equivalent set of hits by up to [Formula: see text] on average versus unguided screening.

Riccardo Curcio, Toni Mancini, Enrico Tronci · 0 citations