Jul 2026
Understanding language model scaling for protein fitness prediction
It is shown that model size, training dataset and stochastic elements can bias the predicted p(sequence) away from real fitness, which clarifies the scaling behavior of protein models on fitness prediction and provides practical guidelines for their application and future development.
Chao Hou, Di Liu, Aziz Zafar et al.
· Nature Computational Science · 2 citations