Skip to content
#protein folding Open access

Ancient Reconstructed Proteins: A Framework for Resurrecting Protein Structures from Million-Year-Old Metagenomes

Aug 2026 · bioRxiv · 0 citations · 67 references
Biology

TL;DR

This work proves that ancient proteins can be reliably recovered from highly degraded palaeogenomic material, establishing a new computational avenue for evolutionary and biochemical research.

Abstract

DNA sequences derived from ancient samples provide insights into human history, paleoenvironments, and evolutionary biology. Advances in laboratory techniques and computational tools have established ancient DNA research as a distinct field. However, current analyses focus mainly on the DNA level, while the protein space remains underexplored. Recent progress in the de novo assembly of ancient metagenomes and the availability of protein structure prediction tools, such as AlphaFold 2, enable the reconstruction of protein structures from these degraded sequences. Here, we present a computational framework to assemble contigs, evaluate their authenticity as ancient sequences, predict open reading frames, and fold ancient protein structures directly from highly damaged metagenomic data. Applying this pipeline to two-million-year-old datasets from the Kap København Formation, we successfully rescued ancient proteins involved in methane metabolism. By generating structural models with AlphaFold 2 and comparing them to modern predicted reference structures, we demonstrate that these ancient proteins can be reconstructed and aligned with high confidence. We showcase this by analyzing an archaeal V/A-type ATP synthase protein recovered from the 2M-year-old Greenlandic data. Ultimately, our work proves that ancient proteins can be reliably recovered from highly degraded palaeogenomic material, establishing a new computational avenue for evolutionary and biochemical research.

Read PDF

Similar papers

Open access Aug 2026

De novo assembly and authentication of ancient DNA metagenomes with nf-core/mag

By introducing support for ancient DNA data in nf-core/mag, this paper aims to improve the ability of researchers to more regularly integrate de novo assembled ancient microbial data into broader metagenomics studies of microbial ecology and evolution.

James A. Fellows Yates, Alexander Hübner, M. Borry et al. · 0 citations
Open access Aug 2026

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline’s output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Guillermo Carrillo-Martin, Johanna Krueger, T. Marquès-Bonet et al. · 0 citations
Open access Jul 2026

Protein Structure Characters in the Light of Phylogenetic Systematics

Abstract Protein structure characters have great potential for improving phylogenetic inference, especially for deep nodes where amino acid sequences are highly diverged. The combination of AlphaFold structure predictions and Foldseek's “3Di” structural alphabet makes it relatively easy to conduct model-based phylogenetic inference that includes a partition of slow-evolving 3Di characters. However, we show that even identical amino acid sequences can produce substantially different 3Di characters, depending on the source of the structural model and whether inter-chain interactions are considered. We argue that such variability can be addressed with key concepts from traditional organism-based phylogenetic systematics: semaphoront, hypodigm, and character ascertainment method. To illustrate this, we develop an analogy between organismal development, taphonomy, and subsequent description and character coding by a systematist, and the process of protein synthesis, folding, and interaction and subsequent extraction, experimentation, and structural modeling by a biochemist. We conclude that differences in 3Di characters between semaphoronts are not intrinsically a problem, but they do require that the researcher uses the same replicable method on all proteins in the phylogenetic analysis. The guiding principle should be to maximize the chance that character differences in the data matrix are the results of underlying evolutionary changes, rather than artifacts due to differences in the methods used for obtaining semaphoronts and coding characters.

Nicholas J. Matzke, Chang-Hao Li · 0 citations
Open access Aug 2026

pastForward: a Snakemake pipeline for ancient and historical DNA with eukaryote-wide taxonomic screening and tracking of copy-number variation

Ancient and historical DNA has the potential to resolve many open questions in biology. While pipelines for processing ancient and historical DNA exist, none combine user-friendly, configurable processing with copy number variation tracking and targeted taxonomic profiling. Therefore, we developed pastForward, a fully automated Snakemake pipeline that integrates all analysis steps from raw reads to damage-rescaled BAM files in a single reproducible workflow. It performs ancient and historical DNA processing, including adapter trimming, read merging, deduplication, damage assessment, quality rescaling, and generates interactive reports summarizing the endogenous read content, library complexity, and breadth and depth coverage statistics. These reports allow users to rapidly assess the quality of sequencing data. It handles single- and paired-end NGS libraries. Mapping to multiple reference sequences is supported, facilitating co-analysis of host and endosymbiont sequences and genotyping of marker genes such as COI. pastForward further integrates two novel tools. ECMSD (Efficient Comprehensive Mitochondrial Sequence Detector) screens each library for eukaryotic DNA by aligning reads against a mitochondrial reference database. The presence of bacteria, archaea and viruses is detected in parallel with Centrifuge. REVEAL (Read-based Estimation and visualization of Element Abundance and Loci) quantifies and visualizes copy number variation of genetic features, such as transposable elements (TEs) or gene duplications. Two case studies demonstrate the usage of the pipeline. Using pastForward on dog genomic time series, including Neolithic samples, we confirm that the copy number of AMY2B, which encodes the starch-digesting enzyme amylase, increased during domestication. From historical D. melanogaster genomes, we recover the recent invasion of the transposable element opus. It is absent in specimens from the 1800s and present from 1933 onward. By efficiently processing large numbers of samples, pastForward facilitates longitudinal tracking of genomic features in diverse species.

Sarah Saadain, M. Kapun, R. Kofler · 0 citations
Open access Aug 2026

An automated pipeline for reconstructing whole genome duplications

Whole genome duplications leave lasting traces in our genomes. How these present in terms of gene content and order varies over time. While collinear blocks of paralogs, long stretches of conserved gene order and content termed ‘microsynteny’, are a distinctive feature of comparatively recent WGD and have been integral in reconstructing the history of ancestral duplication events, this signal degrades over time, making analysis of older events non-trivial. While gene order degrades quickly, gene content is often better conserved and recent work takes advantage of this to reconstruct older events and ancestral pre-WGD and post-WGD chromosomes. However, these new methods are complicated and not well-documented. Here we develop an automated and user-friendly pipeline for reconstructing ancestral chromosomes before and after WGD, and use the conservation of gene content to infer chromosomal rearrangement events in this timeframe. We verify the efficacy of our tool by reconstructing the ancestral acipenseriform, a model system for vertebrate WGD and rediploidisation. Our pipeline should serve to make ancestral reconstruction more accessible and provide a solid foundation for future analysis.

Łukasz Niezabitowski, Anthony K. Redmond, A. McLysaght · 0 citations
Review Open access Aug 2026

Methodological advances and computational frameworks in ancient DNA research: a narrative review

Ancient DNA research has progressed from a limited technical field to a genome-scale discipline capable of reconstructing evolutionary history, past environments, and host-pathogen interactions over long periods. Early studies faced challenges such as severe molecular fragmentation, chemical damage, low endogenous DNA content, and a high risk of modern contamination. Improvements in laboratory techniques and computational methods have enhanced the accuracy and scope of analyses involving degraded DNA. This review synthesizes recent methodological and computational advances across the complete ancient DNA workflow. We discuss how environmental factors, including temperature, pH, and burial context, influence molecular preservation and inform substrate selection, with an emphasis on mineralized tissues and sediments. Key laboratory advances include optimized cleanroom practices, demineralization and preprocessing strategies, silica-based and single-stranded extraction methods for ultrashort fragments, and damage-aware library preparation that balances sequencing efficiency and authenticity. We outline the core authentication criteria, including fragment length distribution, nucleotide misincorporation patterns, endogenous DNA assessment, and contamination monitoring via controls and replication. In addition, we summarize current sequencing strategies and bioinformatic pipelines tailored for short, error-prone reads, highlighting approaches for mapping, contamination estimation, reference bias mitigation, and probabilistic genotype calling in low-coverage datasets. The review further highlights how integrated laboratory and computational frameworks collectively improve the recovery, authentication, and interpretation of ancient genetic material. By integrating methodological rigor with technological innovation, ancient DNA research now enables robust reconstruction of degraded genomes. Continued refinement of laboratory standards and computational frameworks is essential to ensure accuracy, reproducibility, and responsible application in evolutionary, archaeological, and forensic investigations. Future advances in sequencing technologies, authentication strategies, and bioinformatic methodologies are expected to further expand the analytical power and reliability of paleogenomic research. • Ancient DNA research has evolved from PCR-based analysis of short fragments to genome-scale paleogenomic investigations enabled by high-throughput sequencing and specialized laboratory workflows. • Successful recovery of ancient DNA depends on strategic sample selection, with mineralized tissues such as petrous bones and teeth generally providing higher endogenous DNA yields than most other substrates. • Rigorous contamination control, including dedicated cleanroom facilities, demineralization procedures, optimized extraction methods, and damage-aware library preparation, is essential for generating authentic ancient DNA datasets. • Authentication of ancient DNA relies on characteristic molecular signatures, including short fragment lengths, cytosine deamination patterns, endogenous DNA assessment, contamination monitoring, and independent replication. • High-throughput sequencing platforms, particularly Illumina-based technologies combined with target enrichment approaches, have substantially improved the recovery of degraded genomes from low-abundance ancient DNA samples. • Modern bioinformatic frameworks incorporating damage assessment, contamination estimation, reference-bias mitigation, probabilistic genotype calling, and genome reconstruction have become indispensable for accurate interpretation of ancient DNA data.

Julius Luvanga, Abang Akwo, K. Ketu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.