Skip to content

Category

protein folding

562 papers

#protein folding Open access Aug 2026

Title: Bayesian Geometric Chaos with Adaptive Constraint Propagation

Bayesian Geometric Chaos with Adaptive Constraint Propagation represents a novel approach to modeling complex systems, particularly those exhibiting intricate dynamics and high-dimensional parameter spaces. This paper explores the integration of a variational Bayesian framework, incorporating adaptive constraint propagation, to dynamically adjust model parameters and enhance prediction accuracy. Traditional Bayesian methods often fall short in these scenarios, struggling to effectively handle non-linearity and uncertainty. Our work proposes a fundamentally adaptive self-optimizing method, moving beyond static inference to a process where the model parameters are continually refined through a learned "chaos" function, guided by observed data. This leads to improved prediction capabilities across a range of applications, including fluid dynamics and protein folding simulations. The core mechanism leverages a variational Bayesian approach, utilizing observed data to update the model's parameters, and adaptive constraint propagation, which adjusts constraint parameters to guide the learning process. We demonstrate the efficacy of this framework through a series of simulations and analysis, highlighting its potential for addressing limitations of existing Bayesian methods.

Jincheng Zhang · 0 citations
#protein folding Open access Aug 2026

Integrated computational analysis prioritizes candidate targets and pathways linking ochratoxin A exposure to hepatocellular carcinoma

Ochratoxin A (OTA), a food-borne mycotoxin, has been implicated in hepatotoxicity and potential carcinogenic processes, yet the molecular links between OTA exposure and hepatocellular carcinoma (HCC) remain incompletely understood. This study used an integrated computational workflow to prioritize candidate targets and pathways potentially linking OTA exposure with HCC. OTA-related and HCC-related targets were collected from public databases, intersected, and subjected to functional enrichment analysis. Transcriptomic data from the GSE36376 discovery dataset were analyzed to identify differentially expressed genes, followed by LASSO and SVM-RFE feature selection, immune-cell deconvolution, molecular docking, and molecular dynamics simulation. A total of 214 overlapping OTA-HCC-associated targets were identified and were enriched in pathways related to signal transduction, apoptosis, metabolism, and immune regulation. In GSE36376, 443 differentially expressed genes were identified using p < 0.05 and |log2 fold change| > 1, and overlap analysis yielded 13 shared target genes. Five candidate targets, CYP3A4, KIFC1, AKR1C3, CA2, and TTR, were further prioritized. KIFC1 and AKR1C3 were upregulated in HCC samples, whereas CYP3A4, CA2, and TTR were downregulated. These genes showed apparent discriminatory ability within the discovery dataset, with AUC values ranging from 0.866 to 0.958. Molecular docking predicted favorable OTA-target interactions, with docking energies ranging from −7.4 to −10.8 kcal/mol. CYP3A4 showed the lowest predicted docking energy (−10.8 kcal/mol) and was further evaluated by molecular dynamics simulation, with a protein-fitted OTA RMSD of 1.435 ± 0.097 nm and complex Rg of 2.308 ± 0.010 nm during the equilibrated 20–100 ns trajectory. Overall, this study provides a reproducible hypothesis-generating framework for exploring potential metabolic, genomic-instability-related, and immune-microenvironment links between OTA exposure and HCC. Future validation in independent datasets and experimental models will be important to further assess the biological relevance of these candidate targets and pathways.

Shili Yang, Huaiquan Liu, Haiyang Kou et al. · 0 citations
#protein folding Open access Aug 2026

Accurate and efficient prediction of protein conformations with ProtMonomer

Deep learning-based protein structure prediction methods that leverage evolutionary information from multiple sequence alignments (MSAs), exemplified by AlphaFold2, have achieved remarkable accuracy. However, existing methods still struggle to predict challenging proteins, particularly those with novel folds or limited evolutionary information, and to recover alternative conformational states. Here we show that structure prediction models trained under different MSA-depth distributions corresponding to different levels of evolutionary information exhibit complementary generalization behaviors, and that a model trained on a mixture of these distributions can combine their complementary generalization strengths. Building on this insight, we developed ProtMonomer, a deep learning framework trained on MSA-depth distributions representing a broad range of evolutionary information levels to improve structure prediction. Across benchmarks comprising CASP15 targets, non-redundant experimentally determined structures, orphan proteins, and short peptides, ProtMonomer performed comparably to or better than leading methods, including AlphaFold2 and AlphaFold3, with particularly strong performance on challenging targets. For fold-switching proteins, ProtMonomer also recovered alternative conformational states more accurately than AlphaFold2 and AlphaFold3 across diverse homologous sequence sampling strategies. In addition to improving predictive accuracy, ProtMonomer substantially reduced inference cost through an efficient architecture, enabling high-throughput applications. Together, these findings provide insights into the generalization of evolution-informed structure prediction models and support ProtMonomer as an accurate and efficient framework for protein structure prediction.

Yunda Si, Suqi Zhang, Luo-Nan Chen · 0 citations
#protein folding Open access Aug 2026

yvanrousset/FCKcat: FCKcat v1.0.0 – manuscript version

This release contains the version of the FCKcat code associated with the manuscript: Overcoming systematic data biases enables accurate prediction of enzyme kcat fold-changes for computational protein design Authors: Yvan Rousset, Alexander Kroll, and Martin J. Lercher The data and trained model files required to reproduce the analyses are available separately on Zenodo: https://doi.org/10.5281/zenodo.20325541

Yvan Rousset · 0 citations
#protein folding Open access Aug 2026

Three-Dimensional Structural Characterization and Spatial Conformational Ensemble Analysis of the Ultra-Large Multivalent Fusion Protein Construct KH-002v003 (2,091 Amino Acid Residues)

Engineering extended macromolecular therapeutics requires comprehensive structural modeling to verify tertiary folding fidelity and domain accessibility across repetitive structural units. In this study, we present the structural characterization of KH-002v003, an ultra-large synthetic multivalent fusion protein construct expanding to 2,091 amino acid residues. Building upon earlier design iterations—including the 701 aa baseline framework and the 1,354–1,455 aa KH-002v002 architecture—this maximized construct integrates multiple variable heavy-chain nanobody (VHH) domains, tumor microenvironment-cleavable matrix metalloproteinase (MMP-2/9) linkers, pH-low insertion peptides (pHLIP), and C-terminal XTEN solubilization polymers. Structural predictions were executed via high-throughput homology modeling on SWISS-MODEL utilizing a 58-template ensemble superposition. Model 15, constructed against the Cryo-EM structure of the bispecific Fab-heavy chain complex (PDB ID: 8WGW.1.B, sequence identity 60.71%), yielded a peak global QMEANDisCo score of 0.58 ± 0.07. Superposition analysis revealed a dense, rigid central core dominated by antiparallel β-sheet frameworks flanked by dynamic, highly flexible loop regions. Stereochemical validation via MolProbity confirmed 92.16% of residues within favored Ramachandran regions. These findings confirm that ultra-large constructs exceeding 2,000 residues can maintain structural integrity and spatial independence for target engagement.

Khiem Le · 0 citations
#protein folding Open access Aug 2026

Phosphorylation Protects Oncogenic RAS from LZTR1-Mediated Degradation.

Oncogenic KRAS and NRAS mutations are common in hematologic malignancies, but their signaling in this context remains less well characterized than in carcinomas. Using multi-omics screens in multiple myeloma, we sought to identify regulators of RAS activity. We found that the phosphatase PP1C dephosphorylated conserved RAS residue T148, permitting LZTR1-dependent proteasomal degradation. LZTR1 was ineffective against KRAS A146 gain-of-function mutations, which lie adjacent to T148 and are enriched in hematologic cancers, such as diffuse large B cell lymphoma and acute myeloid leukemia. Remarkably, KRAS protein stability was four-fold lower in hematologic versus carcinoma cells, revealing a unique therapeutic opportunity targeting RAS protein stability. PAK1 and PAK2 shielded RAS from LZTR1-dependent degradation by phosphorylating T148, and inhibiting PAK1/2 activity improved RAS-directed therapy. Collectively, these findings reveal a regulatory circuit governing RAS stability that is preferentially active in blood cancers and potentially druggable.

Lin Zhang, A. Bolomsky, Omar S. Al-Odat et al. · 0 citations
#protein folding Open access Sep 2026

V3 Titin Physics — How the Largest Known Protein (34,350 Amino Acids) Folds in 1 Millisecond: A Phase Transition Resolution of the Levinthal Paradox (Ada/SPARK GNATprove 100%)

This Ada/SPARK program applies the V3 Architecture to the largest known protein — titin (connectin) — with 34,350 amino acids. The V3 model resolves the Levinthal paradox by replacing stochastic exploration with a deterministic phase transition guided by four invariants: Ψ_V3 = 48,016.8 kg·m⁻², Φ_critical = -51.1 mV, k = 7 (heptadic closure), and Modulo-9 = 9. The program demonstrates that a protein of 34,350 amino acids folds in 1 ms when phase coherence exceeds 90%. The code is formally verified with GNATprove (100% proof obligations satisfied). It shows that life is a consequence of phase coherence, not random exploration.

outail benhadid · 0 citations
#protein folding Open access Sep 2026

V3 Titin Physics — How the Largest Known Protein (34,350 Amino Acids) Folds in 1 Millisecond: A Phase Transition Resolution of the Levinthal Paradox (Ada/SPARK GNATprove 100%)

This Ada/SPARK program applies the V3 Architecture to the largest known protein — titin (connectin) — with 34,350 amino acids. The V3 model resolves the Levinthal paradox by replacing stochastic exploration with a deterministic phase transition guided by four invariants: Ψ_V3 = 48,016.8 kg·m⁻², Φ_critical = -51.1 mV, k = 7 (heptadic closure), and Modulo-9 = 9. The program demonstrates that a protein of 34,350 amino acids folds in 1 ms when phase coherence exceeds 90%. The code is formally verified with GNATprove (100% proof obligations satisfied). It shows that life is a consequence of phase coherence, not random exploration.

outail benhadid · 0 citations
#diffusion models Dataset Open access Aug 2026

Assessing State-Specific Accuracy of Cofolding Models for Kinases and GPCRs

# Cofolding benchmark: structures, alignments and analysis scripts Supplementary data for *Benchmarking Protein–Ligand Cofolding Models: Correct LigandPlacement Is Decoupled from Accurate Protein Conformation*. Seven protein–ligand complexes, three class A GPCRs and four kinase systems, were predictedwith AlphaFold3, Boltz-2, Chai-1 and RoseTTAFold3 under different MSA and template settingsand compared with the experimental structures. Everything needed to repeat that comparison ishere: the experimental references, the predicted structures, the state-specific alignmentsused to bias the predictions, and the analysis scripts. Tables and figures are not deposited, only the inputs they are made from. The superposedstructures are included, since every reported RMSD was measured on them. ## Contents ```0_reference_structures/ experimental structure of each complex, and the ligand SMILES1_Structures/ the predicted complexes, per target and method2_input_preparation/ scripts that build the model inputs; MSAs/ holds the alignments3_predictions/ how each model was run4_alignment/ superposition onto the reference, and the superposed structures5_ligand_rmsd/ ligand RMSD in the protein reference frame6_extract_metrics/ RMSD, pLDDT, ipTM and PAE collected into per-target tables7_core_analysis_plots/ Figure 48_rmsd_scatter_plots/ Figure 5, Figure S4, and the correlations quoted in the text9_plddt_plots/ Figures S1 and S210_multiseed/ Figure S3, the five-seed repetition11_interaction_fingerprints/ Figures S5-S10, interaction fingerprints and pose validity``` ### Reference structures `0_reference_structures/` holds one folder per target with the experimental complex as mmCIFand as the PDB file the analysis reads, together with the ligand SMILES. The entries are 8ZMG(5-HT2A), 8UGW (A2AAR), 8Y45 (DOR), 9L04 (ALK2 with RK-783), 6UNQ (ALK2 with AMPPNP), 9DMI(LRRK2) and 8TSD (PI3K). ### Predicted structures `1_Structures/` is arranged as target, method, condition: ```1_Structures/GPCR_5HT2a/Alphafold/5ht2a_inactive_custom_templates_custommsa/ seed-0_sample-0/ ... seed-0_sample-4/ one prediction each, .cif and confidences``` The same layout holds for `GPCR_AA2A`, `GPCR_DOR`, `Kinase_ALK2_ATP`, `Kinase_ALK2_RK783`,`Kinase_LRRK2`, `Kinase_PI3K`, and for `boltz`, `Chai1` and `RF3`. Every condition contributesfive diffusion samples of one seed. These are the predictions as the models wrote them; thesuperposed copies live in `4_alignment/superposed_structures/` and, for the five-seed set, in`10_multiseed/superposed_structures/`. ### Biasing alignments `2_input_preparation/MSAs/` holds one alignment per target, named for the conformational stateit biases towards: | file | target | sequences ||---|---|---|| `5HT2A_inactive.a3m` | 5-HT2A | 208 || `A2AAR_active.a3m` | A2AAR | 301 || `DOR_active.a3m` | DOR | 292 || `ALK2_active.a3m` | ALK2, both ligands | 1000 || `LRRK2_dfgout.a3m` | LRRK2 | 669 || `PI3K_dfgout.a3m` | PI3K | 2 | The custom templates are public PDB entries, listed by accession code in Table S2 of theSupporting Information. The scripts beside the alignments turn them into the input files eachmodel expects: AlphaFold3 job JSONs (`create_alphafold3_jobs_new_a3m.py`, and`create_aa2a_active_json.py` for A2AAR), template mmCIFs with the index arrays AlphaFold3needs (`fix_templates_bio.py`), Chai-1 template M8 files (`create_m8_gpcr_chai1.py`), RF3 jobsderived from the AlphaFold3 ones (`create_rf3_jsons_from_af3.py`) and the A3M rewriting RF3requires (`normalize_a3m_for_rf3.py`). ## Requirements `requirements.txt` lists the python packages. The analysis also needs PyMOL for thesuperposition and MDAnalysis for the ligand RMSD; folder 11 additionally needs ProLIF, RDKit,PoseBusters, Open Babel, PDBFixer and OpenMM. The cofolding models themselves are only neededto repeat the predictions, not to reproduce the analysis. ## How to run it Folders 4 to 11 find their inputs inside this deposit, so no paths need editing. Each stepwrites what the next one reads, so run them in this order. Steps 4 and 5 rebuild thesuperpositions and RMSD logs; to work from the deposited superpositions instead, start atstep 6. ```bash# 4-5: superpose each prediction onto its reference, then measure the ligandpython 4_alignment/alignment_pymol.py --savepython 5_ligand_rmsd/lig_rmsd.py # 6: collect RMSD, pLDDT and the confidence metrics into the per-target tablespython 6_extract_metrics/extract_pdb_data.py # the three GPCRspython 6_extract_metrics/extract_kinase_data.py # LRRK2 and PI3Kpython 6_extract_metrics/extract_alk2_data.py # ALK2 with ATP and with RK-783python 6_extract_metrics/extract_confidence.py # ipTM, pTM, interface PAE, overall PAE # 7: Figure 4, and the per-model tables that steps 8 and 9 read.# AXIS_BREAK=3.0 gives the axis break of the published figure; the GPCR panels of# Figure 4 additionally colour best and worst by ligand RMSD.AXIS_BREAK=3.0 python 7_core_analysis_plots/plot_core_analysis_subpart_final.pyAXIS_BREAK=3.0 BEST_BY_LIGAND_RMSD=1 PLOT_OUTDIR=plots_ligbest \ python 7_core_analysis_plots/plot_core_analysis_subpart_final.py # 8: Figure 5, Figure S4, and the correlations quoted in the textpython 8_rmsd_scatter_plots/plot_template_vs_ligand_rmsd.py \ --after-dir 6_extract_metrics/result_tables/after-plotting \ --output-dir plots/scatter --rmsd-cutoff 2.0python 8_rmsd_scatter_plots/plot_plddt_vs_protein_rmsd.pypython 8_rmsd_scatter_plots/plot_plddt_combined.pypython 8_rmsd_scatter_plots/correlation_table.py # 9: Figures S1 and S2python 9_plddt_plots/plot_atomwise_plddt_from_after_csv.py \ --after-dir 6_extract_metrics/result_tables/after-plotting # 10: Figure S3, the five-seed repetitionpython 10_multiseed/aggregate_from_derived.pypython 10_multiseed/create_comparison_heatmaps.py # 11: Figures S5-S10python 11_interaction_fingerprints/prolif_fingerprints.pypython 11_interaction_fingerprints/plot_interaction_heatmaps.py``` Figures land in `plots/`, the multi-seed ones in `10_multiseed/plots/`, and the tables in`6_extract_metrics/result_tables/` and `10_multiseed/result_tables/`. Before step 10, set ` ` at the top of `aggregate_from_derived.py` to the PyMOLexecutable. Without it the script still writes its tables, but the protein RMSD column staysempty. ## How the numbers are defined Predictions are superposed onto the reference on backbone atoms (N, CA, C, O) of a fixedselection, with PyMOL's iterative outlier rejection switched off, so the RMSD covers the wholeselection. For the GPCRs and LRRK2 that selection is the folded core, which leaves out theflexible termini. The ligand RMSD is then measured in that same frame, without superposing theligands onto one another, so it reports where the ligand sits rather than how similar itsinternal geometry is. The extraction scripts add the mean protein and ligand pLDDT from themodel files, and read ipTM, pTM and PAE out of each method's own output. Everything after thatis plotting. ## Steps that were commands rather than scripts The alignments in `2_input_preparation/MSAs/` were built from FASTA sets of structures sharingone conformational state: ```bashmmseqs createdb mmseqs createdb mmseqs search --max-seqs 1000 --threads 8mmseqs result2msa --msa-format-mode 5reformat.pl -M first -r a3m a3m ``` RF3 predictions: ```bashrf3 fold inference_engine=rf3 inputs= ckpt_path= \ diffusion_batch_size=5 seed=0``` Boltz-2 predictions, where the MSA route depends on the condition: `--use_msa_server` for thedefault runs, an explicit A3M for the biased ones, and no MSA at all for the third: ```bashboltz predict --use_msa_server --diffusion_samples 5``` ## Placeholders The scripts in folders 2 and 3 read and write wherever the predictions are run, so they carryplaceholders in angle brackets that need to be filled in. Each script names the ones it usesin a comment under its docstring. | placeholder | what belongs there ||---|---|| ` ` | directory the predictions are written to || ` `, ` `, ` ` | AlphaFold3 job inputs, outputs and model weights || ` `, ` ` | AlphaFold3 sequence databases and the GPU to use || ` `, ` ` | working directory and conda installation of the run scripts || ` ` | PyMOL executable, in `10_multiseed/aggregate_from_derived.py` || ` ` | interpreter of the environment with ProLIF, RDKit and PoseBusters | ## Relation to the earlier version The alignment, ligand RMSD, metric extraction and plotting scripts were revised during peerreview: the superposition runs without iterative outlier rejection on fixed atom selections,and the ligand RMSD is measured in the protein reference frame without refitting. EveryRMSD-derived number and figure in the paper comes from these versions, and the earliervariants are not part of this deposit.

Leon Obendorf, Niklas Piet Doering, Petra Knaus et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling

Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.

Yifan Feng, Guang Cheng, Shihui Ying et al. · 0 citations
#protein folding Sep 2026

Analysis of the Emulsifying Properties of Water-soluble Myofibrillar Proteins Based on Non-covalent Modification with High-charge Density Polysaccharides

This study investigated the effects of non-covalent modification of polysaccharides with different charge density on the emulsifying properties of myofibrillar proteins (MP). Carboxymethyl cellulose (CMC) with degree of substitution values of 0.7, 0.9, and 1.2 (correspond to low, medium, and high charge density) was selected to construct protein-polysaccharide complexes with salt-free MP. The emulsification behavior, rheological characteristics, and stability of CMC-modified MP emulsions were systematically investigated. Results indicated that CMC significantly enhanced the emulsifying activity of MP (P<0.05), with an emulsion activity index reaching 188.53%, representing a 3.83-fold increase. Furthermore, CMC also induced emulsion droplets to exhibit reduced size and narrower distribution. Rheological analysis revealed that high charge density CMC enhanced the elastic and strain-resistant properties of the emulsions. Meanwhile, it strengthened emulsion stability by increasing the apparent viscosity and critical strain levels while inhibiting protein aggregation. This study offers a scientific approach to enhancing the emulsifying properties of MP in low-sodium systems. This approach lays a foundation for the development of low-salt emulsion-based meat products and promoting the application of polysaccharides in novel low-salt meat formulations.

Jingjie WANG, Jiale CHAI, Xinglian XU et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.