MixInfoFold: A Method For Iteratively Generating Better Sequences
Abstract Current AlphaFold3-based protein structure prediction has reached remarkable levels, but sequences that can be applied in practice are mostly generated using physics-based methods. Here, we introduce a graph representation-based protein sequence generation method: MixInfoFold which incorporates new noise addition and feature extraction methods, as well as our designed encoder-decoder called MixInfo. The noise addition method adds Gaussian noise positively correlated with the sequence length at the end of the main-chain, which enhances the model’s ability to understand the dynamic changes in protein structure. The feature extraction method includes a new feature calculated by assessing the area and distance of the main-chain atoms. The MixInfo iteratively computes the relationships between edge and node features, enabling a deeper exploration of the implicit sequence information within protein structures. Compared to the best model, our model achieves an improvement of 0.80%, 0.89%, and 3.28%in sequence recovery rates on the CATH4.2, CATH4.3, and Modelfinal datasets, respectively. Additionally, our model generates sequences faster than other methods, and generates sequences closely resemble natural proteins, indicating the model’s feasibility and potential value in practical applications.