RNA molecules explore heterogeneous conformational ensembles that are essential for their biological function and molecular recognition, yet this intrinsic flexibility poses a major challenge for structure-based drug discovery. In particular, the absence of well-defined binding pockets in static structures limits the identification of ligandable sites. Here, we present an integrative ensemble-based approach that combines enhanced-sampling molecular dynamics simulations with Nuclear Magnetic Resonance data to characterize the conformational landscape of the HIV-1 TAR RNA at atomic resolution. Starting from extensive sampling, we refined the resulting conformational distribution through maximum-entropy reweighting to achieve quantitative agreement with experimental data. Analysis of the reweighted ensemble reveals a diverse set of conformational substates, including compact arrangements that exhibit pocket features compatible with ligand recognition and overlap with known ligand-bound structures. At the same time, highly ligandable conformations, which are only marginally populated, might nonetheless be critical for RNA recognition. Our results demonstrate that integrative ensemble modeling can reveal pharmacologically relevant RNA conformations that are not apparent from experimental static structures, providing a framework for ensemble-based strategies in RNA-targeted drug discovery.
Accurate determination of RNA conformational ensembles is essential for understanding RNA function and advancing RNA-targeted drug discovery, yet lowly-populated alternative states remain difficult to resolve with atomistic detail. A central constraint is that experimental refinement can only select conformations already present in the starting library, making library generation the limiting step. Using the HIV-1 trans-activation response element (TAR) as a model system, we benchmarked conventional MD (cMD) against the enhanced sampling methods Gaussian-accelerated MD (GaMD), replica-exchange Gaussian-accelerated MD (Rex-GaMD), replica-exchange with solute tempering (REST2), and temperature replica-exchange MD (T-REMD), as well as the structure-prediction based methods FARFAR2 and AlphaFold 3. Each library was refined against experimental residual dipolar couplings (RDC) and validated independently using ensemble-averaged QM/MM chemical shifts. We showed that T-REMD produced the most accurate ensemble by both measures, and its advantage tracked with broader, more continuous coverage of the interhelical conformational landscape. Broad temperature-range T-REMD also sampled conformations resembling excited state 1 (ES1) and the U23-A27-U38 base-triple, without requiring these states to be specified during library generation. More accurate ensembles further improved coverage of experimentally observed ligand-bound TAR conformations and enhanced ensemble-based virtual screening, linking structural accuracy to functional utility. The same workflow applied to the preQ1 class I riboswitch and the UUCG tetraloop improved agreement with experimental data in both cases. Together, these results establish replica-exchange enhanced sampling, particularly T-REMD, as an effective strategy for constructing experimentally validated RNA ensembles and accessing conformations corresponding to rare functional substates. Graphical Abstract
This work assesses ML potentials for exploring RNA conformations using the adenine–adenine dinucleoside monophosphate (ApA) dimer, a fundamental RNA building block, and parametrized ML potentials based on the equivariant MACE architecture and informed by both ab initio and semiempirical property data.
Leonardo Medrano Sandonas, Macarena Tolmos Nehme, L. F. Cofas-Vargas et al.· Journal of Chemical Theory a...· 0 citations
Evaluating modern AMBER-based parametrizations across diverse structural motifs, including aptamers, duplexes, and quadruplex-duplex hybrids, provides critical insights for developing next-generation DNA force fields capable of accurately modeling non-native structures and enabling balanced sampling essential for predicting ligand binding in diverse biological contexts.
G. Bekker, Y. Fukunishi, Junichi Higo et al.· Journal of Chemical Theory a...· 0 citations
Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
Hassan Nadeem, D. Kleiman, Yuming Zhou et al.· bioRxiv· 0 citations
Fluorescent base analogs, such as 2-aminopurine (2AP), are frequently used to investigate nucleic acid conformational dynamics, particularly base-stacking interactions, through fluorescence spectroscopy. Although 2AP substitution can induce structural perturbations, properly accounting for these effects enables its use as a probe of complex RNA dynamics. Here, we apply explainable machine-learning-derived surrogate-model collective variables to perform enhanced-sampling simulations that comprehensively explore the free-energy landscapes of both 2AP-substituted and unsubstituted RNA dinucleotides and trinucleotides. Our results show that while 2AP substitution does not significantly alter stacking propensity, it substantially modulates the free energy landscape, leading to a redistribution of populations among stacking modes. The fraction of stacked 2AP conformations closely correlates with experimentally observed “dark state” populations, supporting the hypothesis that base stacking quenches 2AP fluorescence and promotes dark state formation. This work and follow-up studies in larger, physiologically relevant RNA systems will establish the ability to correlate observed emissive or dark states of 2AP with specific structures within the free-energy landscape.
Revanth Elangovan, Emily Stetson, J. Widom et al.· Physical Chemistry, Chemical...· 0 citations
Molecular docking is increasingly used to infer aptamer–target interactions, yet most studies rely on computationally predicted aptamer structures rather than experimentally determined ones. Using a benchmark set of aptamers with known high‐resolution structures, we show that commonly used modeling approaches, including RNAComposer and AlphaFold3, fail to reliably reproduce aptamer conformations, particularly at the binding sites critical for molecular recognition. Key limitations include the use of A‐form RNA models to represent B‐form DNA structures, the prediction of ligand‐free rather than ligand‐bound conformations, and the scarcity of experimentally determined aptamer structures for training machine‐learning models. Using the theophylline aptamer, for which high‐resolution structures are available in both DNA and RNA forms, we systematically evaluated each step of the standard docking workflow. We found that structure‐prediction errors generate incorrect binding pockets, docking scores fail to distinguish theophylline from caffeine despite a 250,000‐fold difference in affinity, and molecular dynamics simulations do not overcome these shortcomings. Together, these results reveal fundamental weaknesses in current aptamer docking workflows and caution against using docking‐derived models to infer binding mechanisms in the absence of experimental structural data.