Foundation models and taxonomy inference for metagenomics: A practical review.
Abstract
Taxonomic classification in metagenomics remains anchored in alignment-based methods and exact k-mer indexers, which provide calibrated, traceable species-level assignments at scale when reference coverage is comprehensive. This review synthesizes the progression of the field toward self-supervised, DNA-specific foundation models, clarifying where they add value and how they should be evaluated. We report a structured literature search with explicit inclusion and exclusion criteria, and then critically compare families of approaches along two axes of practical relevance: reference coverage (in-index versus open-set) and read context and quality (short and accurate versus long and noisy). Across the evidence base, classical k-mer pipelines remain preferable for routine, high-throughput species-level assignment on well-covered clades, whereas the gains reported for foundation models at the read level are mixed and highly sensitive to evaluation design. The most plausible benefits arise under conditions of novelty, for long or noisy sequences, or when a calibrated back-off to higher taxonomic ranks is acceptable; embeddings can also support sample-level augmentation and quality control. Because pretraining and long-context inference impose substantial GPU and memory requirements, practical deployments favor hybrid designs: fast, reference-based classification for the bulk of reads, combined with compact, parameter-efficient adapters or selective embedding to rescore ambiguous cases and flag off-index content. We conclude with a roadmap that emphasizes standardized open-set benchmarks, joint reporting of accuracy and computational cost, and lightweight adaptation techniques that bring foundation-model components within reach of resource-constrained laboratories.