This work constructs genome-free Iso-Seq reference transcriptomes for two organs, flower and leaf, of the non-model species Moricandia arvensis (L.) DC and systematically compares pipeline strategies combining Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction, finding that the optimal pipeline is organ specific.
Abstract
De novo transcriptome assembly is the standard approach for constructing a reference transcriptome in non-model plants that lack a high-quality genome, yet short-read assemblies struggle to resolve full-length isoforms. Long-read Iso-Seq (PacBio) captures full-length transcripts directly, but its use as a primary reference and the choice of downstream assembly pipeline remains poorly benchmarked. Here, we construct genome-free Iso-Seq reference transcriptomes for two organs, flower and leaf, of the non-model species Moricandia arvensis (L.) DC. (Brassicaceae), and systematically compare pipeline strategies combining Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction, benchmarked by BUSCO completeness, RSEM short-read mapping, and TransDecoder ORF completeness. We find that the optimal pipeline is organ specific. For flower, CD-HIT pre-filtering followed by Cogent reconstruction produced a high-quality reference (95.3% BUSCO complete). For the leaf, the same Cogent step was detrimental, reducing BUSCO completeness from 90.1% to 78.0% by incorrectly merging distinct genes; therefore, CD-HIT at 95% identity without reconstruction was retained. We trace this divergence to organ-specific input-data characteristics: leaf transcripts show extreme full-length-read expression skew and predominantly single-isoform gene support, depriving Cogent’s graph algorithm of the multi-isoform evidence it requires. We find that the concentration of full-length reads among the most highly expressed transcripts predicts pipeline suitability before reconstruction, with per-transcript read depth acting as a necessary but non-discriminating floor. Because the leaf reference lacked gene-level structure, we further recovered gene-isoform grouping using expression-aware read-clustering (Corset), which preserved completeness while restoring the paralog structure expected of a paleopolyploid genome and outperformed sequence-only clustering. Our results provide a robust, genome-free framework for constructing full-length reference transcriptomes in non-model plant species and demonstrate that pipeline choice must be evaluated per organ rather than assuming one size fits all.
This study provides a systematic assessment of four commercially available workflows for performing LR snRNA-seq and highlights key methodological trade-offs related to distinct library preparation strategies, thus providing practical guidance for future isoform-resolved transcriptome studies at the single-nucleus leve...
F. Köhler, Anna Delgado-Tejedor, Maik Zehnsdorf et al.· bioRxiv· 0 citations
While isoform identification from long-read RNA sequencing (lrRNA-seq) data has received significant attention, the handling of biologically replicated lrRNA-seq datasets remains less explored. This study defines two strategies for obtaining consensus transcriptomes from multi-sample lrRNA-seq data: Join & Call, where...
Fabian Jetzinger, Alejandro Paniagua, Stanley Cormack et al.· Nature Communications· 1 citation
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome...
Christopher D. R. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al.· bioRxiv· 0 citations
Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms.