Skip to content

Category

protein folding

537 papers

#protein folding Dataset Open access Sep 2026

In silico analysis of SH3BP2 genomic alterations and expression profiles in CRC

Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. The cancer hallmark tool helps in understanding SH3BP2 hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change = 1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-κB and affecting immune responses through WNT/β-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.

Muralidharan Jothimani, Karthikeyan Muthusamy · 0 citations
#protein folding Open access Sep 2026

Non-Linear Constraint Optimization for Biological Systems

This research investigates the application of non-linear constraint optimization to biological systems, specifically focusing on optimization of gene expression and protein folding. Traditional optimization methods often struggle with the inherent complexity and dynamic nature of biological systems, necessitating novel approaches that can effectively capture evolutionary principles and adapt to changing conditions. This paper proposes a hybrid method integrating evolutionary algorithms with constraint optimization, incorporating feedback from the system's own adaptive behavior. We demonstrate the effectiveness of this approach in optimizing a simplified biological model, highlighting its potential for broader applicability in the analysis and control of complex biological processes.

Jincheng Zhang · 0 citations
#protein folding Open access Sep 2026

Title: Dynamical Symmetry Field Analysis

Dynamic Symmetry Field Analysis (DSFA) is a novel technique for unraveling the intricate patterns within complex systems, specifically focusing on identifying and quantifying 'dynamic symmetry fields.' This research explores the potential of employing a new mathematical operator to analyze the evolution of these fields, revealing underlying instabilities and potential for change. The core mechanism leverages the concept of a dynamically evolving symmetry function, allowing for the detection of critical points and deviations from expected behavior. This paper details the mathematical framework, initial results, and potential implications of DSFA for a range of applications, including protein folding, fluid dynamics, and complex mechanical systems. The investigation emphasizes a shift from passive observation to active detection of dynamic structural features, offering a potentially transformative approach to understanding the behavior of such systems.

Jincheng Zhang · 0 citations
#protein folding Dataset Open access Sep 2026

Quadrupling the protein family space with global metagenomics

Here you can find the data from the study "Quadrupling the protein family space with global metagenomics". For Protein Families: iso_clusters25_names.tsv.bz2 Description: TSV file of isolate clusters with more than 25 members. Columns: (1)Cluster name (2)Representative isolate protein header (3) Isolate protein member header (4)Isolate protein member sequence metag_clusters25_names.tsv.bz2 Description: TSV file of metagenomic clusters with more than 25 members. Columns: (1) Cluster name (2) Representative metagenomic protein header (3) Metagenomic protein member header (4) Metagenomic protein member sequence For Protein Folds structures.tar.gz Description: Contains three subfolders with protein structure models HQ:High-Quality (pTM ≥ 0.7) models (26,767 pdb files) MQ:Medium-Quality (0.5 ≤ pTM < 0.7) models(47,823 pdb files) LQ:Low-Quality (pTM < 0.5) models (82,018 pdb files) foldseek_results.tar.gz Description: You will find three folders (HQ, MQ & LQ). Each one contains three files: AF2.tblout (unfiltered hits to AlphaFoldDB) CATH.tblout (unfiltered hits to CATH) PDB.tblout (unfiltered hits to PDB) NMPFAMSDB2_MODELS_SCORES.txt Description: A txt file with pTM and pLDDT score for each family model. Columns: (1) Family name Column (2) pTm score Column (3)pLDDT score

Eleni Aplakidou, Fotis A. Baltoumas, Georgios A. Pavlopoulos · 0 citations
#protein folding Dataset Open access Sep 2026

Local Frustration Modulates the Folding Dynamics of a Repeat Protein

Repeat proteins fold through pathways that are strongly shaped by local energetics, making them highly sensitive to mutations. The ankyrin repeat (AR) domain of IκBα is a cooperative folding unit in which the first four repeats (AR1–AR4) are stable, while the last two are destabilized. Despite this simple modular architecture, IκBα follows a complex folding trajectory involving high-energy intermediates. Here, we investigate how sequence variations modulate folding pathways using coarse-grained AWSEM simulations combined with the Energy Landscape Visualization Method (ELViM) and local frustration analysis. We compare the wild type (WT) with two consensus-designed variants: V93L, which accelerates folding, and L131V, which destabilizes and slows down folding of the protein in vitro. Our results show that WT and V93L share a similar folding funnel, though V93L folds more directly into the native state. In contrast, L131V reshapes the landscape by stabilizing non-native kinetic trap with minimally frustrated contacts. Reversing the L131V mutation allows the protein to reach the native conformation, whereas maintaining it confines the protein to misfolded states that can only be escaped at very high temperatures, without productive folding. These findings highlight how subtle sequence changes can tune frustration, modulate kinetic trapping, and control folding efficiency in repeat proteins.

Murilo N. Sanches, María Inés Freiberger, Peter G. Wolynes et al. · 0 citations
#protein folding Open access Sep 2026

基于自适应的分子动力学模拟优化

Molecular dynamics (MD) simulations are a fundamental tool in materials science, biology, and pharmaceutical research, offering insights into molecular behavior and dynamics. However, traditional MD simulations often suffer from limitations in accuracy and efficiency, particularly when dealing with complex, dynamic environments. This paper introduces a novel algorithm, termed Adaptive Molecular Dynamics Optimization (AMDO), designed to address these challenges by dynamically adjusting simulation parameters based on a learned model. AMDO leverages an adaptive learning algorithm to optimize the simulation process, resulting in enhanced accuracy and reduced computational cost. We demonstrate the effectiveness of AMDO through the simulation of a complex protein folding process, showcasing improved convergence and reduced simulation time compared to conventional MD methods. The core mechanism centers on the continuous adaptation of simulation parameters, enabling the simulation to effectively capture the nuances of molecular interactions.

Jincheng Zhang · 0 citations
#protein folding Open access Sep 2026

The future of metal nanostructure optical tweezers

Abstract Optical tweezers using metal nanostructures have leveraged extreme subwavelength focusing to circumvent the diffraction limit, enabling the isolation and label-free sensing of single nanoparticles. Over the past two decades, nanoaperture optical tweezers (NOTs) have matured into a useful tool for biophysical analysis, monitoring the conformational dynamics, binding affinities, and structural mutations of single proteins without perturbing labels or tethers. Driven by a post-machine-learning shift in biophysics toward understanding the sequence-structure-dynamics-function paradigm, NOTs have recently achieved real-time mapping of single-protein energy landscapes. Future improvements will aim to resolve the sub-microsecond protein folding dynamics by improving the signal-to-noise ratio while navigating the impacts of surface interactions and thermophoresis. This work forecasts key technical innovations over the next five years to achieve nanosecond-scale temporal resolution. By transitioning to smaller metal nanostructures (which includes moving away from nanoapertures) and operating at longer near-infrared wavelengths, near-field sensitivity can be maximized while mitigating laser-induced heating. Augmented functionalities—including integrated Raman spectroscopy (for applications like peptide identification and single-cell proteomics), and enantioselective chiral trapping will expand the utility of metal nanostructure optical tweezers. Combined with machine learning models to maximize data extraction from low signal-to-noise environments and train future predictive models on protein dynamics, these advances aim to deliver a robust platform to understand the dynamics of biomolecules.

Reuven Gordon · 0 citations
#protein folding Open access Sep 2026

CHARGE OPTIMIZATION OF HYDROLASE ENZYMES AND THEIR ASSOCIATED BINDING MODULES FOR WASTE POLYMER DEGRADATION

Lignocellulosic biomass is one of the most abundantly available renewable energy sources, but it is entirely underutilized. Composed of a matrix of cellulose and hemicellulose polysaccharides, biochemical deconstruction of this substrate can yield fermentable sugars for biofuel production. Within a typical biorefinery, enzymatic hydrolysis of these polysaccharides is relied upon by Carbohydrate Active enZymes (CAZymes) to cleave the glycosidic linkages between sugar moieties. This technology is significantly hindered though due to the inherent recalcitrance of biomass to enzymatic degradation. This recalcitrance is caused by challenges related to the structural phenolic polymer lignin which limits enzyme accessibility and non-productively binds CAZymes, as well as low overall activity of the soluble enzymes on a highly insoluble, crystalline substrate. Limited accessibility to polysaccharides is typically alleviated by thermochemical pretreatment of biomass, but this does little to overcome the challenges related to CAZyme functionality. Thus, CAZymes and their appended carbohydrate binding modules (CBMs) must be engineered for improved hydrolytic activity. Lignin and cellulose are both found to have slight negative charges on their surfaces, and exploiting these electrostatics using enzyme supercharging is one promising route for improving enzyme activity. Enzyme supercharging refers to the process of mutating several surface exposed amino acid residues to either positive (R, K) or negative (D, E) charged residues to produce high theoretical net charges. Supercharging was applied to both glycosyl hydrolase (GH) CAZymes and their appended CBMs to tune enzyme surfaces to a critical net charge where binding affinity, and ultimately activity, may be optimized. This principle is closely related to the Sabatier principle which plainly states that enzyme turnover is maximized at intermediate strength binding affinity where neither adsorption nor desorption are limiting factors. This design principle was first applied to a GH family 5 endocellulase Cel5A and its native fused family-2a CBM to generate and systematically screen a supercharged library of CBM2a-Cel5A fusions. As a result of supercharging, several supercharged mutants were identified that showed enhanced binding and catalysis with greater than 2-fold improvements in hydrolysis yields. Enzymes with optimized binding affinity additionally exhibit robust thermotolerance in the presence of substrate compared to the native enzyme. Ultimately, strong substrate dependent correlations between net charge and activity were identified, resembling charge related Sabatier optima. Further probes into these effects were performed by applying this supercharging rationale to a GH family 6 exocellulase Cel6B and its native family-2a CBM. Systematic screening identified key differences resulting from supercharging compared to the endocellulase system. Most strikingly, similar Sabatier effects were not identified, yet a key charge engineered enzyme was isolated from the library. Instead of favorable adsorptive properties, this enzyme exhibited altered unfolding pathways between both domains, revealing a 10 °C increase in melt temperature for the catalytic unit. These processive enzymes possess a more complex mechanism for cellulose degradation, and as such, simple affinity modulation was not an adequate predictor of enzyme performance. Exo-endo synergism is a hallmark of efficient biomass utilization, where endo and exo active cellulases work in conjunction to effectively process cellulose. As such, supercharged enzymes of both classes were combined in several ratios to identify how net charge impacts synergistic hydrolysis. In these cases, the enzymes with the highest individual activity were also the best synergistic partners. These findings suggested that electrostatic engineering may provide a broader strategy for modulating enzyme behavior at heterogeneous polymer interfaces beyond lignocellulosic systems alone. To investigate the transferability of these interfacial engineering principles, charge engineered CBMs were applied to poly(ethylene terephthalate) (PET) synthetic polymer hydrolase systems. Several charge engineered CBMs were fused to a Cutinase enzyme that is capable of hydrolyzing ester linkages in the PET backbone. Results found that, while positively supercharged CBMs enhanced PET binding interactions, hydrolysis improvements did not correlate directly with substrate adsorption. Instead, a slightly negative CBM fusion produced substantial improvements in PET depolymerization through enhanced thermostability and catalytic persistence, revealing that enzyme stability, rather than substrate binding, was the dominant limitation governing hydrolysis performance in this system. The synthesis of this work provides a broader framework to the application of protein supercharging as a rational design technique for increasing enzyme activity. Collectively, the findings in this dissertation demonstrate that heterogeneous biocatalysis cannot be reduced to substrate binding interactions alone. Instead, enzyme activity improvement depends on a dynamic balance between substrate engagement, catalytic turnover, enzyme stability, and catalytic persistence at insoluble polymer interfaces. These results establish electrostatic surface engineering as a powerful, but highly context-dependent strategy for modulating enzyme behavior across diverse heterogeneous polymer degradation systems.

Antonio De Chellis · 0 citations
#protein folding Sep 2026

Antagonistic effects of hypersalinity on heat tolerance in a copepod emerge from non-additive protein abundance changes

Non-additive interactions between environmental stressors, where biological responses to simultaneous stressors do not equal the sum of individual-stressor responses, commonly occur across organisms and environments. To enable predictions of organismal resilience to shifting patterns of environmental stressors, it is important to both identify these interactions and document the mechanisms underlying them. The supratidal copepod Tigriopus californicus demonstrates a non-additive, antagonistic pattern of increased heat tolerance when simultaneously exposed to high salinities. We investigated salinity’s chronic and acute effects on heat tolerance in a northern population and quantified responses in protein abundance to hypersalinity and heat stress, both alone and in combination. The overall proteomic response to multiple stressors was non-additive and largely reflected that to high temperature. However, 42% of multi-stressor proteins were absent from either single-stressor response; we refer to these proteins that are only differentially abundant in the multi-stressor scenario as “emergent”. Our results suggest that the increased heat tolerance of T. californicus conferred by hypersalinity may be driven by a combination of these emergent proteins, several proteins induced by hypersalinity in both single- and multi-stressor conditions that may contribute to cross-tolerance, and four proteins with additive abundance patterns (including a small heat shock protein). These candidate proteins play putative roles in several relevant processes including the heat shock response, protein folding, and regulation of metabolism and oxidative stress responses. Our results connect to prior whole-organism findings and highlight promising pathways for future investigation in the context of heat tolerance and multi-stressor interactions.

C. Terry, Maxime Leprêtre, Dietmar Kültz et al. · 0 citations
#protein folding Dataset Open access Sep 2026

JANUS: data and reproducibility artefacts

Derived data, run outputs and figure source data supporting the article "Joint optimisation of amino acid and coding sequence for de novo designed proteins". Includes the inverse-folding marginals for all 862 backbones and the full 230,992-row double-mutant additivity table.

Anees Ahmed Mahaboob Ali, Radhakrishnan Delhibabu, Everette Jacob Remington Nelson · 0 citations
#protein folding Open access Sep 2026

Omicau Multi-Omics Benchmark Suite

Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It tests predictive performance, modality utility, model-capacity effects, missingness handling, null behavior, exploratory external transport, failure handling, and compute cost without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families with operative structured missingness for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. CPTAC UCEC exploratory holdout: transcriptomics and copy number for endometrioid-versus-serous transport assessment across 95 patient-disjoint tumors. Methods and controls Nine fixed methods compare Omicau with an unmasked architecture-matched ablation, matched early and single-modality neural controls, early and single-modality linear controls, weighted late fusion, a missingness-only diagnostic, and calibrated latent partial least squares. A TCGA-UCEC complete-training-feature sensitivity tests outcome-associated technical missingness. Internal cohorts use shared group-aware partitions, training-only preprocessing, five outer folds repeated three times, 5,000 paired group bootstraps, paired DeLong tests for AUROC, 4,999 paired squared-error sign flips for R-squared, and Holm adjustment across five primary matched-capacity contrasts. Effect sizes and intervals are the primary evidence. Because repeated out-of-fold predictions share training sets, internal p-values are conditional on the frozen prediction vectors and are not unconditional population-generalization tests. Ten target permutations per cohort are coarse catastrophic-leakage diagnostics, not formal empirical tail-probability tests. TCGA and CPTAC expression scales are harmonized by a source-declared, target-blind transformation with pooled and matched-feature numerical-domain gates. The CPTAC endpoint is exploratory because it was exercised during predeposit development smoke. Literature-anchored controls remain independent of method ranking. Failed, unfavorable, discordant, non-estimable, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, internal and external partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation while preserving scientific values and frozen randomization streams. Deviation 2 - Comparator and ablation expansion. Matched neural, unmasked, missingness-only, complete-feature, weighted late-fusion, and partial least-squares controls separate fusion value from model capacity, technical missingness, and integration strategy. All settings are fixed before definitive execution. Deviation 3 - Independent external evaluation. A CPTAC UCEC holdout adds 95 patient-disjoint assessment cases. TCGA-UCEC supplies all training and model-selection rows; shared transcriptomic and copy-number features are aligned by unique Entrez identifiers and expression scales are harmonized without using CPTAC outcomes. Deviation 4 - Primary contrast realignment. The primary contrast is Omicau minus the matched early neural control. Prior linear comparisons remain fully reported as contextual estimates and are not substituted for the matched-capacity test. Deviation 5 - External development exposure. The CPTAC endpoint was exercised during predeposit development smoke. External estimates are designated exploratory and are not treated as untouched confirmatory validation. Deviation 6 - External expression-scale correction. Predeposit development smoke exposed incompatible TCGA and CPTAC expression domains. Linear TCGA RSEM values now receive log2(x+1) after negative values are marked missing; CPTAC retains its source-declared log2 scale. Target-blind pooled and matched-feature gates validate compatibility. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 7 - Synthetic missingness application correction. Predeposit audit showed that registered synthetic missingness masks were not reaching model matrices. The masks now alter every synthetic method input exactly as registered. The correction precedes definitive execution, and earlier smoke outputs are not reused.

TUNA BİRGÜN · 0 citations
#protein folding Open access Sep 2026

Non-Linear Constraint Optimization for Biological Systems

This research investigates the application of non-linear constraint optimization to biological systems, specifically focusing on optimization of gene expression and protein folding. Traditional optimization methods often struggle with the inherent complexity and dynamic nature of biological systems, necessitating novel approaches that can effectively capture evolutionary principles and adapt to changing conditions. This paper proposes a hybrid method integrating evolutionary algorithms with constraint optimization, incorporating feedback from the system's own adaptive behavior. We demonstrate the effectiveness of this approach in optimizing a simplified biological model, highlighting its potential for broader applicability in the analysis and control of complex biological processes.

Jincheng Zhang · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.