Aug 2026· Bioinformatics· Vol 42· 0 citations· 33 references
Medicine
TL;DR
Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms.
Abstract
Abstract Motivation Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. Results We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. Availability CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244.
KazRNA-Pipe is presented, an open-source Nextflow v26.6 workflow processing both bulk and single-cell RNA sequencing (RNA-seq) within a unified Singularity-containerized environment with graphics processing unit (GPU) acceleration through the RAPIDS ecosystem.
M. Ashimgaliyev, Beimbet Daribayev, A. Zhumadillayeva et al.· BioMedInformatics· 0 citations
Abstract Motivation Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or...
Abhishek Agarwal, Ziad Al Bkhetan, Dariusz Plewczynski· Bioinformatics· 0 citations
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome...
Christopher D. R. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al.· bioRxiv· 0 citations
SingleCellMQC is an open-source R package that provides a unified QC framework for single-cell RNA sequencing, surface proteome profiling, and immune repertoire data, and its modular architecture allows flexible integration with existing workflows.
Dai-Han Ji, Mei Han, Shu-Ting Lu et al.· iScience· 0 citations
Abstract Motivation Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, resea...