MsaTM-DB: a large-scale empirical database linking alignment properties, substitution models, and gene-tree metrics in phylogenomics
Abstract
Phylogenomic inference requires empirical datasets that capture the diversity of molecular evolutionary processes, including heterogeneity in substitution processes and phylogenetic signal, yet existing resources rarely provide standardized, per-locus alignment, model, and tree metrics. We present MsaTM-DB, a curated database comprising 965,545 loci from 420 eukaryotic phylogenomic studies, each annotated with 35 features spanning alignment properties, substitution parameters, and gene tree metrics. These data enable systematic investigation of heterogeneity in phylogenetic signals across diverse evolutionary contexts. An integrated pipeline and interactive R Shiny platform support distributional analyses, correlation exploration, and empirically informed simulations. Using this database, we illustrate its downstream potential through exploratory analyses showing that alignment- and tree-derived features can help evaluate factors influencing phylogenetic support. MsaTM-DB provides an extensive empirical foundation for benchmarking phylogenetic methods, guiding marker selection, and developing data-driven evolutionary models. Our database is available online at the GitHub repository (https://github.com/xtmtd/MSA-and-tree-metrics-exploration).