The results show that examining branching configurations optimal modulo the branching energy provides structural information beyond the standard MFE prediction, and the proposed energy-filtering approach yields a compact set of alternative structural hypotheses that can complement Boltzmann sampling and provide a practical source of candidate helices for downstream computational or experimental analysis.
Abstract
Accurately predicting RNA secondary structure remains a central challenge in computational biology. Although standard methods based on minimum free energy (MFE) optimization often produce good predictions, their accuracy varies, and they can still miss helices present in known structures. Because the thermodynamics of multibranch loops strongly influence these predictions, we examine how changes to the multiloop initiation and branching penalties affect prediction performance, focusing specifically on recall. Using a recent algorithm that partitions the branching-parameter space, we generate all distinct optimal structures obtainable under different choices of the multiloop parameters. We then evaluate the recall of these alternative structures for the Archive II dataset. Our results show that many sequences admit multiple alternative structures with substantially higher recall than the MFE prediction, establishing the predictive potential of multiloop reparameterization. We next introduce an energy-based filtering method that retains only those structures whose adjusted residual energy is at least as good as that of the MFE structure. This produces a tractable number of candidates while preserving most of the achievable improvements in recall. Compared with Boltzmann sampling, the resulting ensemble typically provides a more favorable balance between recall and precision at the level of helix classes despite being much smaller, making it particularly useful for identifying lower-probability structural features. Overall, our results show that examining branching configurations optimal modulo the branching energy provides structural information beyond the standard MFE prediction. The proposed energy-filtering approach yields a compact set of alternative structural hypotheses that can complement Boltzmann sampling and provide a practical source of candidate helices for downstream computational or experimental analysis.
Non-coding RNAs play diverse roles in a wide range of cellular processes, with their spatial structure being pivotal to their function. RNA secondary structure is a key determinant of its overall fold. Given the scarcity of experimentally determined RNA 3D structures, understanding secondary structure is vital for discerning RNA function. Currently, there is no universally effective solution for de novo RNA secondary structure prediction. Existing methods are becoming increasingly complex without marked improvements in accuracy and often overlook critical features such as pseudoknots and alternative folds. Here, we introduce SQUARNA, a new approach to de novo RNA secondary structure prediction that is suitable for both individual RNA analysis and large-scale structural searches. SQUARNA revisits the concept of base pair maximization and develops it into a stem maximization idea coupled with the widely used free energy minimization (MFE) framework. SQUARNA can predict alternative structures and handle pseudoknots of arbitrary complexity. Benchmarking shows that SQUARNA outperforms existing methods, including deep learning models, in both single-sequence and alignment-based RNA secondary structure prediction. SQUARNA seamlessly integrates sequence and alignment information with experimental data, such as residue reactivities obtained by chemical probing, as well as other structural restraints, including automated searches for Rfam database templates, G-quadruplex patterns, and protein-binding motifs. SQUARNA is available as a standalone tool at https://github.com/febos/SQUARNA and as a web server at https://larnal.imol.institute.
Maksim D. Serdakov, Davyd R. Bohdan, Grigory Nikolaev et al.· bioRxiv· 0 citations
QSAD is presented, a quantum-classical framework that reformulates peptide structure prediction as amino-acid-level Hamiltonian sampling and replaces iterative optimization with non-iterative Hamiltonian evolution and establishes coarse-grained quantum sampling as a practical computational path for structure prediction in regimes where data-driven methods lack sufficient signal.
Yuqi Zhang, Bo Fang, Yuxin Yang et al.· 1 citation
RNA structure is central to the function of every RNA class yet the gap between annotated sequences and experimentally determined structures remains large. Computational methods to fill this gap have evolved from thermodynamic free energy minimization through supervised deep learning to self-supervised RNA language models trained on millions of sequences, progressively improving structure prediction. Here we review the state of the art in RNA structure prediction, covering key training datasets, community benchmarks, and the performance of current models. We further discuss perspectives on integrating other data modalities, such as chemical probing signals and RNA modifications, as well as the emerging role of generative models. Challenges in generalization, handling of noncanonical interactions, and contextual structure prediction remain open frontiers for the field.
Lambert Moyon, Annalisa Marsico· Current Opinion in Structura...· 0 citations
Deep learning has advanced RNA secondary-structure prediction by bypassing explicit energy rules to capture long-range dependencies, yet progress is limited less by model scale than by how structures are measured: single scores hide where and why models fail, and benchmark scores can reflect memorization of one dataset rather than genuine generalization. We address this with Shifu, a framework of three coupled parts. Shifu-Corpus is a leakage-audited dataset of 254123 sequences from six databases, with family-aware splits certified free of exact and near-duplicate leaks. The Shifu Trifecta scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number. Shifu-LMR, a family of compact RNA language models, serves as controlled experiments: changing the training corpus shifts accuracy by 0.13, and a 65-million-parameter model, Shifu-LMR-Nano, leads on correctness while running on a laptop. We release the dataset, code, and model backbones.
Gabriel Galvez, Quentin Vicens· bioRxiv· 0 citations
The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.
Lars A. Eicholt, Lasse Middendorf· bioRxiv· 0 citations