Skip to content
Open access

Handling biological replicates in long-read RNA sequencing data by joining or not joining

Sep 2026 · Nature Communications · Vol 17 · 1 citation · 50 references
Medicine

Abstract

While isoform identification from long-read RNA sequencing (lrRNA-seq) data has received significant attention, the handling of biologically replicated lrRNA-seq datasets remains less explored. This study defines two strategies for obtaining consensus transcriptomes from multi-sample lrRNA-seq data: Join & Call, where reads from all samples are combined before transcript reconstruction, and Call & Join, where transcript reconstruction is performed on individual samples before combining the resulting annotations. We apply these strategies to mouse brain and kidney tissue datasets, using PacBio and ONT technologies, across six transcript reconstruction tools. Our results indicate that the optimal strategy depends on the tool and research objective. We find that Join & Call is generally preferable for discovering novel isoforms, while Call & Join is often preferable for highly replicated datasets when the discovery of novelty is secondary. Our findings provide a conceptual and practical framework for multi-sample transcript reconstruction, guiding best practices for increasingly large-scale lrRNA-seq studies. When combining long-read RNA sequencing data from multiple samples, transcripts can be identified either before or after merging. This study shows that neither approach is universally optimal and provides guidance based on the software tool and study goals.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.