Skip to content

A Scalable Pan-Genomic Pipeline for Annotation-Free Discovery of Species-Specific Markers: Application to Staphylococcus aureus

Jul 2026 · Journal of Bioinformatics and Computational Biology · 0 citations

TL;DR

An open-source, annotation-independent pan-genomic pipeline featuring an overlapping sliding-window algorithm and a three-tier subtractive screening funnel successfully circumvents conventional gene-centric limitations, providing a generalizable computational strategy for target discovery across other high-priority bacterial pathogens.

Abstract

Current bioinformatics approaches for bacterial diagnostic target discovery remain constrained by their reliance on gene annotations and fixed-boundary genome segmentation, which overlook unannotated intergenic regions and introduce sequence-truncation artifacts. Here, we developed an open-source, annotation-independent pan-genomic pipeline featuring an overlapping sliding-window algorithm (500-bp window, 100-bp step) and a three-tier subtractive screening funnel. Using Staphylococcus aureus as a model, the pipeline screened 1,629 genomes against 852 non-S. aureus Staphylococcus genomes and >20,000 background bacterial genomes. Seven highly conserved, unannotated targets (SA-1 to SA-7) were identified, with all seven translated into qPCR primer sets (SAP-1 to SAP-7), among which three (SAP-1 to SAP-3) were further characterized by in vitro experiments. Multi-layer in silico evaluation demonstrated 100% intraspecific sensitivity and zero cross-reactivity against background genomes, including the S. aureus complex. In vitro testing using crude cell lysates confirmed specific amplification of S. aureus DNA without non-target cross-reactivity, establishing a qualitative limit of detection (LOD) of 10 5 CFU/mL. Additional computational validation on draft genomes, raw sequencing reads, near-neighbor species, and a clinical truth set corroborated marker robustness under realistic conditions. This framework successfully circumvents conventional gene-centric limitations, providing a generalizable computational strategy for target discovery across other high-priority bacterial pathogens.

View source

Similar papers

Open access Jul 2026

Pandoomain, a scalable pipeline for genomic and protein domain context analysis, reveals widespread PT-TG domain architectural diversity and novel polymorphic toxins

ABSTRACT The rapid expansion of bacterial genome databases presents significant opportunities for functional discovery, as a large fraction of genes and protein domains remain uncharacterized. Analyzing genomic context and domain architecture is a powerful approach for functional inference, but existing tools often lack the scalability and integrated workflow required for high-throughput analysis. To address this, we developed Pandoomain, a Snakemake pipeline that automates the acquisition of genomes from the National Center for Biotechnology Information, identifies proteins of interest using hidden Markov models (HMMs), and performs systematic domain annotation and gene neighborhood analysis. We demonstrate the utility of Pandoomain through a comprehensive analysis of the poorly characterized pre-toxin TG (PT-TG) domain across 347,289 bacterial genomes. Our analysis revealed 10,226 PT-TG-containing proteins organized into 312 unique domain architectures, highlighting their association with diverse interbacterial antagonistic systems, including the Type VI secretion, Type VII secretion, and contact-dependent inhibition systems. By leveraging genomic context, we identified a novel variant of the WXG trafficking domain, termed W10XG, and subsequently discovered 24 new families of associated toxin domains. We experimentally validated six of these toxins, confirming that all six are neutralized by their cognate immunity proteins. Pandoomain is an accessible tool that enables systematic, large-scale exploration of protein domains, and our analysis of the PT-TG domain provides a rich resource for future investigations into the mechanisms and evolution of bacterial antagonism. IMPORTANCE The rapid growth of bacterial genomic data presents a major hurdle for scientists seeking to understand the functions of newly discovered genes and proteins. To address this issue, we created Pandoomain, a powerful, accessible software tool that automates large-scale analysis of genetic information across hundreds of thousands of genomes. Using Pandoomain, we investigated a poorly understood family of proteins involved in bacterial competition, revealing novel protein domain architectural diversity. This led to the discovery of 24 new families of toxins predicted to be used by bacteria to attack their competitors, and we experimentally confirmed the toxic activity of six of them. Our work provides the scientific community with a robust tool to accelerate functional discovery and offers new insights into the evolution of bacterial conflicts, which may provide insights into the compositional dynamics of microbial communities and support methods to engineer their composition. The rapid growth of bacterial genomic data presents a major hurdle for scientists seeking to understand the functions of newly discovered genes and proteins. To address this issue, we created Pandoomain, a powerful, accessible software tool that automates large-scale analysis of genetic information across hundreds of thousands of genomes. Using Pandoomain, we investigated a poorly understood family of proteins involved in bacterial competition, revealing novel protein domain architectural diversity. This led to the discovery of 24 new families of toxins predicted to be used by bacteria to attack their competitors, and we experimentally confirmed the toxic activity of six of them. Our work provides the scientific community with a robust tool to accelerate functional discovery and offers new insights into the evolution of bacterial conflicts, which may provide insights into the compositional dynamics of microbial communities and support methods to engineer their composition.

E. Soto, Adam Oliver, Marcos H. de Moraes · 0 citations
Review Open access Aug 2026

Variant calling in non-model organisms with snpArcher.

Population genomic studies in non-model organisms increasingly depend on whole-genome resequencing, yet translating raw reads into reliable variant callsets remains a practical challenge due to the complexity of multi-step bioinformatics pipelines and the absence of species-specific best practices. Here we present a step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis. We guide users through six phases: installation and environment setup, sample sheet creation, run configuration, execution on local or high-performance computing systems, quality control review using an interactive HTML dashboard, and downstream analysis, focusing on postprocessing and filtering. The QC dashboard aggregates individual-level metrics including principal component analysis, relatedness estimation, depth-missingness diagnostics, and admixture analysis to help identify batch effects, contamination, cryptic relatedness, and outlier samples before downstream analysis. We demonstrate the impact of sequential filtering steps on the site frequency spectrum and demographic inference using a dataset of 137 burrowing owl (Athene cunicularia) genomes, showing how removal of low-coverage individuals, sex-linked scaffolds, and regions of excess heterozygosity eliminates artifacts that would otherwise bias inference of population size history. This protocol is intended as a practical companion to the original snpArcher publication, enabling researchers working with non-model organisms to produce and evaluate analysis-ready variant callsets in a reproducible manner.

Cade Mirchandani, Abdelmajid Omarjee, Guillaume Achaz et al. · 0 citations
Open access Aug 2026

Applying a proteogenomic approach for improving genome annotation in Leishmania panamensis using high-resolution mass spectrometry data.

Leishmania panamensis, a protozoan parasite of the Viannia subgenus, causes American tegumentary leishmaniasis (ATL) throughout Central and South America, with clinical manifestations ranging from cutaneous lesions to mucosal involvement. Despite the availability of a reference genome for L. panamensis MHOM/COL/81/L13 strain, comprising 30.69 megabases across 35 chromosomes and 8,665 predicted protein-coding genes, comprehensive proteogenomic analyses to refine these annotations have remained limited. This study utilised a proteogenomic approach to enhance the genome annotation of the L. panamensis MHOM/COL/81/L13 strain by integrating publicly available liquid chromatography-tandem mass spectrometry data with a custom six-frame translated genome database. Through systematic peptide mapping and validation, we identified 50 novel protein-coding genes previously absent from the reference annotation and corrected 50 existing genes. The newly identified genes encode proteins containing functional domains, including thioredoxin, glucosyltransferases, and myotubularin, which likely contribute to parasite metabolism, host-pathogen interactions, and intracellular survival mechanisms. These findings substantially refine the genomic resource available for L. panamensis research and provide potential targets for diagnostic and therapeutic development. Future studies will incorporate orthogonal validation methods such as RNA-seq and Ribosome profiling to distinguish between functional genes and translational noise. The refined annotation enables more accurate functional genomic studies and enhances our understanding of the molecular basis underlying L. panamensis pathogenicity and immune evasion strategies. Moreover, the peptide-supported novel and corrected protein-coding genes identified in this study represent a valuable resource for the future exploration of candidate biomarkers and may be provide the molecular targets for improved diagnosis, prognosis, and therapeutic intervention in American tegumentary leishmaniasis.

Soumi Chowdhury, Shubhankar Pawar, Praveen Kumar et al. · 0 citations
Open access Aug 2026

Genome-context-aware discovery of antibacterial peptides from bacterial small open reading frames

Small open reading frames (sORFs) are a potentially rich, yet error-prone, source of antimicrobial-peptide (AMP) candidates: short sequences are readily prioritized by AMP classifiers but may derive from incomplete gene calls. We developed a genome-context-aware discovery workflow that separates AMP-like sequence properties from evidence for a complete, recurrent coding locus. From 649,653 RefSeq assemblies representing 327 clinically relevant bacterial species, species-aware clustering and length filtering yielded 4,442,548 representative 10–100-aa sequences. AmpScanner v2, Macrel and AMPlify identified 585 non-haemolytic records supported by all three models. However, genome-context auditing of 11,918 mapped candidates showed that 529 of 536 mapped consensus candidates were supported exclusively by partial ORFs near contig termini. By contrast, 3,382 candidates had at least one complete non-edge occurrence; 1,069 recurred in ≥2 assemblies and 251 in ≥10 assemblies. We therefore assembled a 20-peptide panel through two explicitly labelled routes: sequence/structure-led selection (n=8) and genome-supported selection (n=12). Broth microdilution against Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 25923 identified low-micromolar activity in both routes. CAND_04141, a recurrent complete non-edge candidate, had the strongest combined profile (MICs of 4 and 2 μM, respectively), while CAND_07825 and CAND_04265 were also active at low micromolar concentrations. In plate-count MBC assays, all three advanced peptides achieved ≥3-log10 reductions at 128 μM. These findings show that high classifier agreement is not a substitute for genomic evidence and provide an auditable framework for prioritizing both synthetic AMP-like sequences and candidate genome-encoded peptides.

Qingxiu Li, Zhenjun Li · 0 citations
Open access Jul 2026

PhytoFam: A Nextflow Pipeline for Genome-Wide Analysis of Plant Gene Families

Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complicates reproducibility and parameter tracking, and may require additional steps to remove redundant protein isoforms. To address these challenges, we developed PhytoFam, a Nextflow-based workflow that automates gene family identification from proteome input through phylogenetic reconstruction. The pipeline integrates HMMER for candidate sequence identification, isoform-aware deduplication, InterProScan for domain confirmation, BLAST reciprocal best hit (RBH) analysis for orthology assignment, MUSCLE for multiple sequence alignment with optional outgroup incorporation, TrimAl for alignment trimming, and IQ-TREE3 for phylogenetic reconstruction. PhytoFam is portable across local workstations and high-performance computing environments and supports deployment through Conda, Docker, and Singularity. We validated the workflow using the Morus alba MADS-box gene family, where the complete analysis finished in 1 h 10 min (9 CPU h). IQ-TREE3 accounted for most of the execution time, whereas InterProScan showed the highest memory requirement with a peak resident set size of 4.5 GB. PhytoFam provides a reproducible, automated, and scalable solution for plant gene family identification and phylogenetic analysis. The pipeline is freely available at https://github.com/sanamparajuli/PhytoFam.

Sanam Parajuli, Bibek Adhikari, Anne Y. Fennell et al. · 0 citations
Open access Jul 2026

Long-read metagenomics and methylation-based binning support the discovery of antibiotic resistance gene-host associations in complex communities.

By linking ARGs to their wider genetic contexts and hosts, the findings shed light on the previously unrecognized carriers of resistance genes in wastewater, and provides a valuable methodology for early identification of newly arising ARGs and their hosts.

Melina A. Markkanen, Heidi Putkuri, D. Kičiatovas et al. · 0 citations