Injecting synthetic competitors drawn from each target’s own wrong-answer score distribution, holding the true site, candidate pool and ranker fixed, reduced top-five recovery by 16.8 points on training folds and 17.0 points on the held-out test fold.
Abstract
Cryptic-pocket prediction is compared almost entirely by top-n recovery, which conflates two separable abilities: proposing a candidate at the true site, and ranking it highly enough to be seen. We retained per-candidate overlaps for five candidate-generation methods, evaluated in six configurations, on the CryptoBench test fold. Coverage spanned 14.5 points; conversion of coverage into top-five recovery spanned 36.7. Two rankers over an identical candidate set differed by 10.6 recovery points at equal coverage, isolating ranking exactly. Union coverage saturated at 92.2%, reaching 98.6% on sites of at least eight residues, with residual failures concentrated on small sites. Injecting synthetic competitors drawn from each target’s own wrong-answer score distribution, holding the true site, candidate pool and ranker fixed, reduced top-five recovery by 16.8 points on training folds and 17.0 points on the held-out test fold. Pooling detectors consequently gains nothing at a budget of five and 11.8 points at twenty.
Pockets from a protein language model at locations where geometry finds no concavity improves single-structure recovery by 8.5% (95% CI +4.0 to +13.6) on test-fold data, and lets a five-conformer ensemble match a twenty-conformer one at a third of the wall clock.
Lacuna, an open-source Python tool for discovering cryptic binding pockets, generates a conformational ensemble from any input structure, detects pockets independently in every conformer, clusters the detections into persistent sites across the ensemble, and ranks those sites with a model fitted on within-structure pairs.
Novo-1, a coarse-grained cofolding framework for binding- affinity prediction, offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, and demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets.
Nikhil Shenoy, David Errington, Emmanuel Bengio et al.· bioRxiv· 0 citations
Alignment-free lineage assignment from k-mer frequency profiles is widely used for SARS-CoV-2 surveillance, and the methods that do it are ranked against each other by margins of one or two percentage points. Those rankings rest on an unchecked protocol. Public repositories hold many near-duplicate genomes, and stratified random splitting puts members of such a group on both sides of the split, so a classifier is credited for sequences it has already seen. We propose quantised profile hashing, which finds near duplicates in k-mer feature space by rounding each frequency vector and hashing it. No sequence is compared with any other, so one pass over the feature matrix suffices and no similarity threshold has to be chosen. Rounding is also what makes the groups well defined, and they are then kept whole across the training, validation and test sets. On 255,611 genomes from seven Pango lineages, random splitting leaves 5.09% of test sequences with a near duplicate in training, on a benchmark ranked by margins of one or two points. Ten update rules were trained twice, identically except for the partition. The contaminated benchmark separates one rule from the leader at 0.05; the clean one separates none. The two orderings are uncorrelated, Kendall τ = +0.022, with rules moving 3.2 positions on average and the leader of one benchmark ranking eighth on the other. A ranking obtained under contamination therefore says nothing about the ranking without it, and the quantity worth reporting beside a score is the leakage rate of the split.
Identifying cryptic binding sites in proteins remains a challenge in structure-based drug discovery because these sites are often not apparent in apo structures. Here, we developed and validated a novel "induce-and-identify" workflow that integrates mixed solvent molecular dynamics (MxMD) simulations with SiteMap. This approach leverages MxMD to sample protein conformations to expose hidden pockets, which are then effectively identified and ranked by SiteMap. Using a challenging data set of 65 cryptic binding sites, the developed workflow identified the cryptic binding site within the top 5 predictions in 78.5% of cases. These results suggest that the proposed MxMD + SiteMap workflow provides a robust and valuable tool for early phase drug discovery, enabling the exploration of a broader range of druggable targets by effectively inducing and identifying cryptic binding sites.
Da Shi, Dmitry Lupyan, Steven V. Jerome et al.· Journal of Chemical Informat...· 0 citations