The GenomeCompendium is released, a public database and interactive analysis tool for complete prokaryotic genomes and it is shown that complex, repeat-rich genomes are more common than previously estimated.
Abstract
Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq (∼47,000) and GenBank (∼13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classification and frequency analysis, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ∼90 features across genomes, offers downloadable reports and -as unique features-pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.
Wastewater treatment relies on complex microbial communities, yet existing genome-resolved references for this essential engineered ecosystem remain dominated by short-read assemblies, limiting genome contiguity and linkage between taxonomic and metabolic function. We applied long-read sequencing to activated sludge from 83 globally distributed plants, reconstructing 53,501 metagenome-assembled genomes to establish the Microbial Database of Activated Sludge (MiDAS) global genome catalog. The catalog encompasses high-quality genomes for 12,047 prokaryotic species, 82% of which are not represented in GTDB release 226, and provides a median of 32 high-quality genomes for each of the 250 core prokaryotic genera previously defined in our MiDAS global 16S rRNA gene survey. This enables analyses of predicted functional traits and their ecological context, for example, we identified two sparsely represented Nitrospiraceae genera with conserved nitrite-oxidation genes that are abundant in higher-temperature wastewater treatment plants. In summary, the MiDAS genome catalog provides a framework for linking taxonomy, metabolism and ecological roles in wastewater treatment systems globally.
Lei Liu, C. Singleton, R. Kirkegaard et al.· bioRxiv· 0 citations
Understanding of the metabolic capabilities and genomic landscape of the P. fluorescens species is enhanced, providing a foundation for natural product discovery using bioinformatic approaches.
Sajid Iqbal, Farida Begum· Discover Genetics and Evolut...· 0 citations
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.
C. D. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al.· bioRxiv· 0 citations
Chloroplast or plastid genomes are essential resources for studying the evolution and diversity of algae and land plants. Although thousands of plastid genomes have been sequenced, their full potential has not been realised; derived resources such as orthogroup databases and reference datasets for metagenomic profiling remain underdeveloped. We present the ChlORIS database to address these problems across all algal phyla. From 2,254 publicly available algal plastid genomes, after dereplication we clustered 2,531 orthogroups from the annotated proteins and selected 496 orthogroups with consistent gene naming, enabling cross-genome comparisons of homologous plastid proteins. We further selected 224 core orthogroups, each containing more than 10 protein sequences, for which we produced score-calibrated hidden Markov models (HMMs), multiple sequence alignments and predicted protein structures. The value of these resources for phylogenomics is demonstrated through a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages. We characterised the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection. The metagenomic database, HMM library, nucleotide and amino acid alignments, predicted structures and protein metadata, cross-linked to UniProt and InterPro (Pfam), are openly available on the ChlORIS website at https://chloris.codeberg.page/.
Yuhao Tong, Vanessa Rossetto Marcelino, Robert Turnbull et al.· bioRxiv· 0 citations