Prediction and Entropy of Evolving Sequences: A Transformer-Based Estimator of Functional Information
Abstract
To resist thermodynamic disorder, living systems maintain low entropy genomic configurations through selection. We propose pseudo-perplexity, the per position surprisal of a masked language model, as an estimator of this biological functional information. Using a unified framework, we demonstrate that transformer derived surprisal systematically distinguishes functionally constrained sequences from unconstrained ones across digital and biological substrates. In evolved populations of 64-byte BFF replicators, a fine-tuned GPT-2 assigns significantly lower pseudo-perplexity to conserved opcode positions than to variable regions (Mann-Whitney U, r = 0.99). A fine-tuned Nucleotide Transformer reproduces this pattern for bacterial 16S rRNA sequences (p = 1 × 10−16, r = 0.2), assigns lower pseudo-perplexity to essential than non-essential genes (p = 6.1 × 1020, r = 0.14), and reveals a negative power law relationship between genomic pseudo-perplexity and essential gene fraction across 12 organisms (Spearman Ï = −0.58, p = 0.048). Finally, pseudo-perplexity of variable regions increases with phylogenetic distance while conserved regions remain relatively stable, recapitulating a diffusion-like drift in both digital and biological lineages. These results establish transformer surprisal as a potential measure of functional information in evolving systems. Data/Code available at: https://github.com/tydymy/bff-functional-information