Skip to content
Open access

DyProL: Dynamic Ensemble Representation Learning for Protein–Nucleic Acid Binding Site Prediction

Sep 2026 · Advancement of science · 0 citations · 53 references
Medicine

Abstract

ABSTRACT Protein–nucleic acid interactions play central roles in gene regulation and cellular function, and extensive efforts have been devoted to predicting nucleic acid binding sites from protein structures. However, protein–nucleic acid recognition is inherently dynamic, whereas most existing computational approaches rely on single static conformations, limiting their ability to capture conformational heterogeneity underlying binding. Here, we present DyProL, an ensemble‐based conformational representation learning framework that models proteins as ensembles of conformations sampled from equilibrium‐like structural distributions. DyProL learns dynamic structural features through iterative aggregation of intra‐ and inter‐conformation geometric information, enabling representation of both local structural context and global conformational variability. Across multiple benchmarks, DyProL consistently outperforms state‐of‐the‐art methods in nucleic acid binding site prediction, with particularly pronounced improvements under realistic settings using predicted or apo‐like structures, where static methods degrade substantially. These results establish dynamic ensemble‐based representations as a general and scalable paradigm for structure‐based protein modeling, providing a foundation for improving a broad range of protein function prediction tasks.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.