Jun 2026· ACM International Conference on Bioinformatics, Computational Biology and Biomedicine· Vol 2026, pp. 1-9· 0 citations· 29 references
Computer ScienceMedicine
TL;DR
UniPocket is presented, a unified multitask framework for residue-level prediction of both ligand-binding and cryptic-pocket residues within a single shared-backbone architecture and achieves a macro ROC-AUC of 0.76.
Abstract
Identifying druggable pockets in proteins is central to structure-based drug discovery, yet conventional ligand-binding-site prediction and cryptic-pocket prediction are typically treated as separate tasks requiring distinct tools, preprocessing pipelines, and evaluation protocols. This separation is limiting because the two problems share a common biological foundation: co-evolutionary patterns encoded in protein sequences carry implicit signals about both conventional and cryptic binding sites. We present UniPocket, a unified multitask framework for residue-level prediction of both ligand-binding and cryptic-pocket residues within a single shared-backbone architecture. UniPocket uses frozen per-residue ESM-2 embeddings as input to a lightweight residual MLP with two task-specific heads, trained on ligand-contact labels from recently deposited PDB structures and cryptic-pocket annotations from CryptoBench. Training alternates between ligand-labeled and cryptic-labeled mini-batches, while an orthogonality regularizer encourages the two heads to learn complementary rather than redundant signals. UniPocket achieves a macro ROC-AUC of 0.82 for cryptic-pocket prediction and 0.76 for ligand-binding-site prediction — matching or exceeding all dedicated single-task baselines on both tasks simultaneously, without any 3D structural input at inference time.
Predicting small molecule-protein interactions across nonhomologous proteins remains challenging because shared ligand recognition is often not evident from sequence, fold, or pocket similarity. Here, we introduce pocket hopping, a machine-learning framework that learns residue-level interaction patterns from coligand binding pockets and infers compatibility between nonhomologous pockets for similar chemotypes. Using shared ligands as supervision rather than explicit geometric alignment, pocket hopping identifies pocket relationships that are not readily captured by conventional chemical-, sequence-, or structure-based comparisons. In two case studies, pocket hopping demonstrates broad utility in drug discovery by enabling de novo hit identification and mechanistic interpretation, identifying fedratinib and its analogues as helicase WRN inhibitors. The model also identified the clinical-stage HDAC inhibitor abexinostat as a direct ENPP1 binder and inhibitor, and cellular assays showed enhanced cGAMP-STING signaling under cGAMP stimulation. Together, these results indicate that pocket-level compatibility can complement existing approaches for target identification, hit discovery, and polypharmacology analysis.
Yingying Zhang, Tianbiao Yang, Buying Niu et al.· Journal of Medicinal Chemist...· 0 citations
A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.
Ryan Varghese, Pooja Tiwary, Krishil Oswal· bioRxiv· 0 citations
A pipeline reformulating kinase-substrate modeling as a Bayesian inference problem is presented and it is revealed that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores.
Jinyuan Hu, Shimian Li, Yue Xue et al.· Journal of Chemical Informat...· 0 citations
The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.
Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.
P. Bayat, Spencer J. Perkins, Sebastian Clancy et al.· bioRxiv· 0 citations
Accurate protein family classification is essential since proteins within the same family share conserved structural domains and biochemical functions that deter mine their biological roles. G protein-coupled receptors (GPCRs) represent one of the largest and most diverse protein families in eukaryotes, serving as targets for ap proximately 35% of FDA-approved drugs. While traditional sequence alignment methods, such as BLAST, provide foundational tools for identifying homologous sequences, they exhibit limited accuracy in distinguishing closely related GPCR families with low sequence homology. Recently, deep learning approaches offer promising accuracy; however, they employ fixed-size classification architectures that force newly discovered protein families into pre-existing categories, preventing the recognition of novel families and limiting scalability as the protein universe expands. In this work, we present a scalable machine learning framework GPCR SLM, that classifies GPCRs across 86 distinct families using a lightweight transformer model optimized through knowledge distillation. Our approach achieved an overall ac curacy of 99%, significantly outperforming BLAST (86.1%) and HMMER (91%), while demonstrating substantial computational efficiency with an average speedup of 33.5× compared to large protein language models. These results demonstrate the effectiveness of combining distilled protein language models with flexible classification frameworks for high-resolution functional annotation.
Fairuz Shadmani Shishir· IEEE transactions on computa...· 0 citations