The results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.
Abstract
Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.
Long-read RNA sequencing directly resolves the full structures of RNA transcripts. Advances in throughput now enable the generation of deeply sequenced cohorts of hundreds of samples, making joint transcript discovery across large datasets possible. However, existing transcript identification methods were designed for...
E. Dolzhenko, Megan D. Schertzer, Ryan Gossart et al.· bioRxiv· 1 citation
This study provides a systematic assessment of four commercially available workflows for performing LR snRNA-seq and highlights key methodological trade-offs related to distinct library preparation strategies, thus providing practical guidance for future isoform-resolved transcriptome studies at the single-nucleus leve...
F. Köhler, Anna Delgado-Tejedor, Maik Zehnsdorf et al.· bioRxiv· 0 citations
Cell-to-cell transcriptional heterogeneity, or noise, is an intrinsic property of the transcriptome with implications for development, disease progression, and aging. Bulk RNA-seq masks this variability by averaging gene expression across cells, whereas single-cell RNA sequencing (scRNA-seq) resolves it. Nevertheless,...
Rafal Czapiewski, M. Chiang, James Ding et al.· bioRxiv· 0 citations
A single genomic assay that delivers complete information across variant classes remains an aspirational goal. Currently, researchers and clinicians rely on an inefficient, expensive combination of short-read sequencing for single-nucleotide variants (SNVs) and small indels, comparative genomic hybridization (CGH) ar...
Single-cell RNA-sequencing resolves cellular states in exquisite detail. Yet oncogenic gene fusions, key drivers in 16.5% of malignancies and ∼50-70% of acute lymphoblastic leukaemia (ALL) cases, remain largely invisible at this resolution. This leaves a fundamental gap in understanding cancer biology. We close it with...
Jovana Maksimovic, Victoria Streeton-Cook, Calandra V. Grima et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.