BACKGROUND
Protein sequence and structure similarity-based search is an important task, which underpins protein annotation, evolutionary analysis, large-scale functional inference, and the exploration of the protein "dark space". The rapid growth of sequence and predicted structure databases has spurred diverse search methods, yet their evaluation remains limited to fold-level similarity and inconsistent benchmarking protocols.
RESULTS
We present a comprehensive benchmark for protein sequence and structure search. Using this framework, we evaluate 14 representative methods spanning sequence alignment, structure alignment, and representation-based approaches across multiple biologically relevant scenarios. Our results show pronounced and context-dependent differences among methods. Structure alignment methods excel at detecting fold-level and geometric similarity, while representation-based searching approaches show advantages in capturing functional similarity under low sequence identity and robustness to predicted structures. Notably, all evaluated methods show limited effectiveness on intrinsically disordered proteins.
CONCLUSIONS
This benchmark establishes a standardized framework for evaluating protein similarity search methods, providing a practical resource for method selection and a foundation for the development of next-generation approaches capable of addressing diverse homology search challenges.
Yuan Liu, Yingquan Zhou, Yan Huang et al.· Genome Biology· 1 citation
Accurate immune receptor design requires modeling the coupled variation of aminoacid sequence, full-atom conformation, and target-binding geometry across antibodies, nanobodies, and T-cell receptors (TCRs). Existing methods often address only part of this problem, either by separating structure generation from sequence design, relying on fixed-backbone inverse folding, or focusing on a single receptor class. We introduce IgGM2, a unified all-atom generative framework for immune receptor structure prediction and CDR sequence–structure co-design. IgGM2 follows a structure-to-design strategy: it first learns how immune receptors are positioned around fixed target structures, and then transfers this target-conditioned structural prior to CDR design. Unlike modular design pipelines, IgGM2 jointly generates CDR residue identities and full-atom receptor structures, allowing frame-work geometry to adapt to designed CDRs without separate inverse folding or external sidechain packing. Unlike continuous residue encodings based on virtualatom geometry, IgGM2 keeps sequence prediction explicit while using atom14 placeholders only for full-atom representation. On structure prediction benchmarks, IgGM2 better captures receptor–target spatial relationships than AlphaFold3 on FoldBench and achieves strong performance on TCR–pMHC modeling. On sequence design benchmarks, IgGM2 achieves competitive amino-acid recovery and improves Rosetta-based interface preference metrics, suggesting more favorable generated binding interfaces. These results support IgGM2 as a unified all-atom framework for adaptive immune receptor structure prediction and design.
Jian Ma, Fandi Wu, Lin Yao et al.· bioRxiv· 0 citations