Skip to content
Open access

Benchmarking long-read RNA sequencing for de novo transcriptome assembly in non-model plant species: insights from Moricandia arvensis

Sep 2026 · bioRxiv · 0 citations · 38 references
Biology

TL;DR

This work constructs genome-free Iso-Seq reference transcriptomes for two organs, flower and leaf, of the non-model species Moricandia arvensis (L.) DC and systematically compares pipeline strategies combining Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction, finding that the optimal pipeline is organ specific.

Abstract

De novo transcriptome assembly is the standard approach for constructing a reference transcriptome in non-model plants that lack a high-quality genome, yet short-read assemblies struggle to resolve full-length isoforms. Long-read Iso-Seq (PacBio) captures full-length transcripts directly, but its use as a primary reference and the choice of downstream assembly pipeline remains poorly benchmarked. Here, we construct genome-free Iso-Seq reference transcriptomes for two organs, flower and leaf, of the non-model species Moricandia arvensis (L.) DC. (Brassicaceae), and systematically compare pipeline strategies combining Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction, benchmarked by BUSCO completeness, RSEM short-read mapping, and TransDecoder ORF completeness. We find that the optimal pipeline is organ specific. For flower, CD-HIT pre-filtering followed by Cogent reconstruction produced a high-quality reference (95.3% BUSCO complete). For the leaf, the same Cogent step was detrimental, reducing BUSCO completeness from 90.1% to 78.0% by incorrectly merging distinct genes; therefore, CD-HIT at 95% identity without reconstruction was retained. We trace this divergence to organ-specific input-data characteristics: leaf transcripts show extreme full-length-read expression skew and predominantly single-isoform gene support, depriving Cogent’s graph algorithm of the multi-isoform evidence it requires. We find that the concentration of full-length reads among the most highly expressed transcripts predicts pipeline suitability before reconstruction, with per-transcript read depth acting as a necessary but non-discriminating floor. Because the leaf reference lacked gene-level structure, we further recovered gene-isoform grouping using expression-aware read-clustering (Corset), which preserved completeness while restoring the paralog structure expected of a paleopolyploid genome and outperformed sequence-only clustering. Our results provide a robust, genome-free framework for constructing full-length reference transcriptomes in non-model plant species and demonstrate that pipeline choice must be evaluated per organ rather than assuming one size fits all.

Read PDF

Similar papers

Open access Sep 2026

Systematic benchmarking of commercial workflows for isoform-resolved single-nucleus transcriptomics

This study provides a systematic assessment of four commercially available workflows for performing LR snRNA-seq and highlights key methodological trade-offs related to distinct library preparation strategies, thus providing practical guidance for future isoform-resolved transcriptome studies at the single-nucleus leve...

F. Köhler, Anna Delgado-Tejedor, Maik Zehnsdorf et al. · 0 citations
Open access Sep 2026

Handling biological replicates in long-read RNA sequencing data by joining or not joining

While isoform identification from long-read RNA sequencing (lrRNA-seq) data has received significant attention, the handling of biologically replicated lrRNA-seq datasets remains less explored. This study defines two strategies for obtaining consensus transcriptomes from multi-sample lrRNA-seq data: Join & Call, where...

Fabian Jetzinger, Alejandro Paniagua, Stanley Cormack et al. · 1 citation
Review Open access Oct 2026

Deciphering transcriptome complexity via long‐read sequencing

Abstract Transcriptomics is moving beyond gene‐level quantification toward isoform‐resolved interrogation of alternative splicing, transcript structural variation, and repeat‐derived transcription. Yet short‐read sequencing remains intrinsically limited in accurately reconstructing full‐length transcripts and resolving...

Chu-Wen Xu, Jia Li, Chen-Xi Yin et al. · 0 citations
Open access Aug 2026

nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality

The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome...

Christopher D. R. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.