Aug 2026· Frontiers in Plant Science· Vol 17· 0 citations· 29 references
Medicine
TL;DR
This work presents REAPER (Repeatome Extended Analysis Pipeline—Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses.
Abstract
Introduction Repeatome characterization from short-read sequencing data is widely performed using RepeatExplorer2/TAREAN. However, long-lived multisample projects and explicit comparative designs are often executed as ad hoc command sequences that are hard to version, rerun, and monitor on shared compute environments — a gap that motivates a project-centric workflow layer for repeatome analysis. Methods We present REAPER (Repeatome Extended Analysis Pipeline—Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses. REAPER does not implement a new repeat-discovery algorithm; it is an orchestration layer, and biological accuracy for clustering and satellite calling depends on the underlying RepeatExplorer2/TAREAN and satMiner methods it coordinates. REAPER standardizes: Read QC Deterministic subsampling and preparation RepeatExplorer2/TAREAN execution via seqclust, with satMiner-inspired iterative assembly Post-TAREAN BLAST-based annotation against curated repeat collections (optionally including taxon-scoped NCBI-derived resources with freshness checks) Optional graph-based comparative reports The pipeline makes comparative read allocation, prefix policy, and analysis-ready tables explicit; caching supports incremental reruns and structured logs support monitoring. Performance was assessed using a Triticeae short-read dataset (five samples), with rule-level logging of runtime and memory across pipeline stages. Results Rule-level performance logs show that graph-based clustering dominates runtime and memory, while QC and preparation steps are lightweight by comparison. Graph-report annotations for the Triticeae project additionally link high-ranking clusters to established repeat markers — including pTa794- and pSc119-class entries in curated databases. Discussion These findings illustrate biologically interpretable outputs (recovery of known Triticeae repeat markers) alongside quantitative performance metrics (identification of graph-based clustering as the dominant computational cost). By making comparative read allocation, prefix policy, and analysis-ready tables explicit — and by supporting caching and structured logging — REAPER supports reproducible comparative repeatome analysis in evolving multisample projects. As an orchestration layer rather than a discovery algorithm, REAPER's contribution lies in reproducibility, monitorability, and comparative-analysis infrastructure, with biological accuracy remaining contingent on the underlying RepeatExplorer2/TAREAN and satMiner methods.
Summary Current tools for RNA epitranscriptomic modification and structural feature analysis are often fragmented, focusing on single signal types and struggling with processing efficiency, especially as sequencing data volumes increase and single-cell technologies advance. To address this challenge, we developed Modte...
Tong Zhou, Yifan Hong, Panfeng Li et al.· bioRxiv· 1 citation
The results suggest that current coding agents can accelerate scaffolding, documentation, and routine implementation, but they do not eliminate the need for expert review in bioinformatics workflow construction.
Paul R. Munn, Jennifer K. Grenier· bioRxiv· 0 citations
A step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis, is presented.
Cade Mirchandani, Abdelmajid Omarjee, Guillaume Achaz et al.· Molecular biology and evolut...· 0 citations
SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data and generates a merged consensus callset across the analysis cohort, which can be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant pri...
Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms.