Skip to content

In Silico Identification and Comparative Synteny of Biosynthetic Gene Clusters in Plants.

2026 · Methods in molecular biology · Vol 3054, pp. 15-27 · 0 citations
Medicine

TL;DR

A bioinformatics approach combining plantiSMASH and MCScan for the identification, annotation, and comparative analysis of BGCs in plant genomes, which can significantly enhance research in plant genomics and metabolic engineering by offering new insights into the organization and function of BGCs.

View source

Similar papers

Open access Jul 2026

PhytoFam: A Nextflow Pipeline for Genome-Wide Analysis of Plant Gene Families

Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complicates reproducibility and parameter tracking, and may require additional steps to remove redundant protein isoforms. To address these challenges, we developed PhytoFam, a Nextflow-based workflow that automates gene family identification from proteome input through phylogenetic reconstruction. The pipeline integrates HMMER for candidate sequence identification, isoform-aware deduplication, InterProScan for domain confirmation, BLAST reciprocal best hit (RBH) analysis for orthology assignment, MUSCLE for multiple sequence alignment with optional outgroup incorporation, TrimAl for alignment trimming, and IQ-TREE3 for phylogenetic reconstruction. PhytoFam is portable across local workstations and high-performance computing environments and supports deployment through Conda, Docker, and Singularity. We validated the workflow using the Morus alba MADS-box gene family, where the complete analysis finished in 1 h 10 min (9 CPU h). IQ-TREE3 accounted for most of the execution time, whereas InterProScan showed the highest memory requirement with a peak resident set size of 4.5 GB. PhytoFam provides a reproducible, automated, and scalable solution for plant gene family identification and phylogenetic analysis. The pipeline is freely available at https://github.com/sanamparajuli/PhytoFam.

Sanam Parajuli, Bibek Adhikari, Anne Y. Fennell et al. · 0 citations
Aug 2026

Coffea Comparative Genomics Reveals Subgenome-Associated Expansion and Diversification of Biosynthetic Gene Clusters.

Biosynthetic gene clusters (BGCs) are important for plant specialized metabolism, but remain poorly characterized in coffee. Given Brazil's importance in coffee production, we performed a comparative genomic analysis of BGCs across the allotetraploid Coffea arabica and its diploid progenitors C. eugenioides and C. canephora. Using standardized genome filtering, annotation, orthogroup inference, and cluster classification, we identified 472 BGCs comprising 3,118 biosynthetic genes, which were grouped into 194 cluster families and integrated with 28,280 orthogroups. Of these, 7,091 orthogroups were shared across all species; Coffea canephora and subgenomes shared 10,923, while Coffea eugenioides and subgenomes shared 11,446. Most BGC-associated orthogroups (86.4%) link to a single pathway class. BGC-associated genes form a highly structured yet lineage-dynamic component of the Coffea pangenome. C. eugenioides and its derived subgenomes in Arabica contributed disproportionately to 14 BGC-associated orthogroups, including flavonoid-, lipid-, and stress-related functions. In contrast, C. canephora derivatives contributed only two terpene-related orthogroups. The parental species showed fewer secondary metabolism-related enriched GO terms (3 and 1) than their subgenomes (53 and 56). Species-specific rearrangements, expansions, and subgenome retention indicate that hybridization and polyploidy shaped BGC diversification. These results advance understanding of specialized metabolism in Coffea and identify targets for coffee improvement and climate resilience.

D. Chacon, L. Gonzalez-García, Vitor Trinca et al. · 0 citations
Open access Jul 2026

BGX: A Comprehensive Pipeline for Genomic Insight into Bioactivity Prediction, Genomic Surveillance, and Novel Biosynthetic Gene Cluster Assessment

The increasing availability of genomic and metagenomic data has created significant opportunities to explore microbial diversity, biosynthetic potential and functional traits. However, comprehensive and comparative genome analysis often requires integrating multiple independent tools, making large-scale studies challenging to implement and manage. Here, we present Bacterial Genome eXplorer (BGX), an integrated and scalable pipeline that streamlines large-scale genome analysis and exploration of biosynthetic potential. BGX integrates different analytical tools into six major stages: (i) genome retrieval and assembly, (ii) genome quality assessment, (iii) antimicrobial resistance (AMR) gene profiling, (iv) annotating Biosynthetic gene clusters (BGCs) and novelty assessment, (v) bioactivity predictions, and (vi) clustering and networking analysis. In addition, BGX provides a user-friendly interactive interface to facilitate data exploration and interpretation. We demonstrated the versatility and scalability of BGX through large-scale analysis of two independent datasets: 248 genomes from the One Day One Genome (ODOG) initiative and 153 publicly available genomes from NCBI. This analysis enabled the comprehensive characterisation of genome quality, AMR determinants, biosynthetic potential and candidate bioactive metabolites across two datasets. BGX is distributed as a Docker container that simplifies installation, enables reproducible data processing, and supports pipeline execution across different computational environments. The modular and reproducible architecture of the BGX pipeline provides an effective framework for large-scale genome mining, genomic surveillance, and accelerates the discovery and prioritisation of novel secondary metabolites. BGX is now accessible at https://bgx.nabi.res.in

Ardhendu Chakrabortty, Lovepreet Singh, Babanpreet Kaur et al. · 0 citations
Open access Aug 2026

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae , and Streptomyces griseorubens . Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur , and Nur , which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

Nada S. Al-Theyab, Haila M. Alnassar, Mohanad A. Ibrahim et al. · 0 citations
Open access Jul 2026

Evolution, structure and function of the putative biosynthetic gene cluster of the fungal secondary metabolite myriocin, a potent inhibitory sphingolipid.

This work identifies the putative myriocin biosynthesis gene cluster (BGC) through de novo sequencing of two producing fungi, Isaria sinclairii and Mycelia sterilia, yielding genomes of 27 and 20 secondary metabolite BGCs, respectively.

B. Rutter, Michael A. Herrera, Gustavo Perez Ortiz et al. · 0 citations