Skip to content
Open access

GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes

Aug 2026 · bioRxiv · 0 citations · 81 references
Biology

TL;DR

The GenomeCompendium is released, a public database and interactive analysis tool for complete prokaryotic genomes and it is shown that complex, repeat-rich genomes are more common than previously estimated.

Abstract

Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq (∼47,000) and GenBank (∼13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classification and frequency analysis, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ∼90 features across genomes, offers downloadable reports and -as unique features-pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.

Read PDF

Similar papers

Review Open access Jul 2026

The MiDAS global genome catalog: 53,501 long-read MAGs representing all core prokaryotic genera in the global activated sludge microbiome

Wastewater treatment relies on complex microbial communities, yet existing genome-resolved references for this essential engineered ecosystem remain dominated by short-read assemblies, limiting genome contiguity and linkage between taxonomic and metabolic function. We applied long-read sequencing to activated sludge from 83 globally distributed plants, reconstructing 53,501 metagenome-assembled genomes to establish the Microbial Database of Activated Sludge (MiDAS) global genome catalog. The catalog encompasses high-quality genomes for 12,047 prokaryotic species, 82% of which are not represented in GTDB release 226, and provides a median of 32 high-quality genomes for each of the 250 core prokaryotic genera previously defined in our MiDAS global 16S rRNA gene survey. This enables analyses of predicted functional traits and their ecological context, for example, we identified two sparsely represented Nitrospiraceae genera with conserved nitrite-oxidation genes that are abundant in higher-temperature wastewater treatment plants. In summary, the MiDAS genome catalog provides a framework for linking taxonomy, metabolism and ecological roles in wastewater treatment systems globally.

Lei Liu, C. Singleton, R. Kirkegaard et al. · 0 citations
Open access Aug 2026

nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality

The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.

C. D. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al. · 0 citations
Open access Aug 2026

ChlORIS: Chloroplast Orthologs Resource & Identification Suite

Chloroplast or plastid genomes are essential resources for studying the evolution and diversity of algae and land plants. Although thousands of plastid genomes have been sequenced, their full potential has not been realised; derived resources such as orthogroup databases and reference datasets for metagenomic profiling remain underdeveloped. We present the ChlORIS database to address these problems across all algal phyla. From 2,254 publicly available algal plastid genomes, after dereplication we clustered 2,531 orthogroups from the annotated proteins and selected 496 orthogroups with consistent gene naming, enabling cross-genome comparisons of homologous plastid proteins. We further selected 224 core orthogroups, each containing more than 10 protein sequences, for which we produced score-calibrated hidden Markov models (HMMs), multiple sequence alignments and predicted protein structures. The value of these resources for phylogenomics is demonstrated through a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages. We characterised the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection. The metagenomic database, HMM library, nucleotide and amino acid alignments, predicted structures and protein metadata, cross-linked to UniProt and InterPro (Pfam), are openly available on the ChlORIS website at https://chloris.codeberg.page/.

Yuhao Tong, Vanessa Rossetto Marcelino, Robert Turnbull et al. · 0 citations