2026· Methods in molecular biology· Vol 3054, pp.
15-27
· 0 citations
Medicine
TL;DR
A bioinformatics approach combining plantiSMASH and MCScan for the identification, annotation, and comparative analysis of BGCs in plant genomes, which can significantly enhance research in plant genomics and metabolic engineering by offering new insights into the organization and function of BGCs.
Understanding of the metabolic capabilities and genomic landscape of the P. fluorescens species is enhanced, providing a foundation for natural product discovery using bioinformatic approaches.
Sajid Iqbal, Farida Begum· Discover Genetics and Evolut...· 0 citations
Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complicates reproducibility and parameter tracking, and may require additional steps to remove redundant protein isoforms. To address these challenges, we developed PhytoFam, a Nextflow-based workflow that automates gene family identification from proteome input through phylogenetic reconstruction. The pipeline integrates HMMER for candidate sequence identification, isoform-aware deduplication, InterProScan for domain confirmation, BLAST reciprocal best hit (RBH) analysis for orthology assignment, MUSCLE for multiple sequence alignment with optional outgroup incorporation, TrimAl for alignment trimming, and IQ-TREE3 for phylogenetic reconstruction. PhytoFam is portable across local workstations and high-performance computing environments and supports deployment through Conda, Docker, and Singularity. We validated the workflow using the Morus alba MADS-box gene family, where the complete analysis finished in 1 h 10 min (9 CPU h). IQ-TREE3 accounted for most of the execution time, whereas InterProScan showed the highest memory requirement with a peak resident set size of 4.5 GB. PhytoFam provides a reproducible, automated, and scalable solution for plant gene family identification and phylogenetic analysis. The pipeline is freely available at https://github.com/sanamparajuli/PhytoFam.
Sanam Parajuli, Bibek Adhikari, Anne Y. Fennell et al.· bioRxiv· 0 citations
Biosynthetic gene clusters (BGCs) are important for plant specialized metabolism, but remain poorly characterized in coffee. Given Brazil's importance in coffee production, we performed a comparative genomic analysis of BGCs across the allotetraploid Coffea arabica and its diploid progenitors C. eugenioides and C. canephora. Using standardized genome filtering, annotation, orthogroup inference, and cluster classification, we identified 472 BGCs comprising 3,118 biosynthetic genes, which were grouped into 194 cluster families and integrated with 28,280 orthogroups. Of these, 7,091 orthogroups were shared across all species; Coffea canephora and subgenomes shared 10,923, while Coffea eugenioides and subgenomes shared 11,446. Most BGC-associated orthogroups (86.4%) link to a single pathway class. BGC-associated genes form a highly structured yet lineage-dynamic component of the Coffea pangenome. C. eugenioides and its derived subgenomes in Arabica contributed disproportionately to 14 BGC-associated orthogroups, including flavonoid-, lipid-, and stress-related functions. In contrast, C. canephora derivatives contributed only two terpene-related orthogroups. The parental species showed fewer secondary metabolism-related enriched GO terms (3 and 1) than their subgenomes (53 and 56). Species-specific rearrangements, expansions, and subgenome retention indicate that hybridization and polyploidy shaped BGC diversification. These results advance understanding of specialized metabolism in Coffea and identify targets for coffee improvement and climate resilience.
D. Chacon, L. Gonzalez-García, Vitor Trinca et al.· Genome· 0 citations
The increasing availability of genomic and metagenomic data has created significant opportunities to explore microbial diversity, biosynthetic potential and functional traits. However, comprehensive and comparative genome analysis often requires integrating multiple independent tools, making large-scale studies challenging to implement and manage. Here, we present Bacterial Genome eXplorer (BGX), an integrated and scalable pipeline that streamlines large-scale genome analysis and exploration of biosynthetic potential. BGX integrates different analytical tools into six major stages: (i) genome retrieval and assembly, (ii) genome quality assessment, (iii) antimicrobial resistance (AMR) gene profiling, (iv) annotating Biosynthetic gene clusters (BGCs) and novelty assessment, (v) bioactivity predictions, and (vi) clustering and networking analysis. In addition, BGX provides a user-friendly interactive interface to facilitate data exploration and interpretation. We demonstrated the versatility and scalability of BGX through large-scale analysis of two independent datasets: 248 genomes from the One Day One Genome (ODOG) initiative and 153 publicly available genomes from NCBI. This analysis enabled the comprehensive characterisation of genome quality, AMR determinants, biosynthetic potential and candidate bioactive metabolites across two datasets. BGX is distributed as a Docker container that simplifies installation, enables reproducible data processing, and supports pipeline execution across different computational environments. The modular and reproducible architecture of the BGX pipeline provides an effective framework for large-scale genome mining, genomic surveillance, and accelerates the discovery and prioritisation of novel secondary metabolites. BGX is now accessible at https://bgx.nabi.res.in
The genus
Streptomyces
is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated
Streptomyces
strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as
Streptomyces thinghirensis, Streptomyces novocaesareae
, and
Streptomyces griseorubens
. Applying the consensus framework across the three
Streptomyces
genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems;
Fur, Zur
, and
Nur
, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified
Streptomyces
isolates as a source of novel natural products.
Nada S. Al-Theyab, Haila M. Alnassar, Mohanad A. Ibrahim et al.· Frontiers in Microbiology· 0 citations
This work identifies the putative myriocin biosynthesis gene cluster (BGC) through de novo sequencing of two producing fungi, Isaria sinclairii and Mycelia sterilia, yielding genomes of 27 and 20 secondary metabolite BGCs, respectively.
B. Rutter, Michael A. Herrera, Gustavo Perez Ortiz et al.· Communications Biology· 0 citations