Dynamic Symmetry Field Analysis (DSFA) is a novel technique for unraveling the intricate patterns within complex systems, specifically focusing on identifying and quantifying 'dynamic symmetry fields.' This research explores the potential of employing a new mathematical operator to analyze the evolution of these fields, revealing underlying instabilities and potential for change. The core mechanism leverages the concept of a dynamically evolving symmetry function, allowing for the detection of critical points and deviations from expected behavior. This paper details the mathematical framework, initial results, and potential implications of DSFA for a range of applications, including protein folding, fluid dynamics, and complex mechanical systems. The investigation emphasizes a shift from passive observation to active detection of dynamic structural features, offering a potentially transformative approach to understanding the behavior of such systems.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It evaluates predictive performance, modality utility, null behavior, failure handling, and compute cost without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. Design and controls All real cohorts were fixed after source and endpoint eligibility checks and before method performance was observed. The design uses shared group-aware partitions, training-only preprocessing, five outer folds repeated three times for real cohorts, 40 independent synthetic replicates, ten target permutations per real cohort, 5,000 paired group bootstraps, and Holm adjustment across the five primary real-dataset contrasts. Literature-anchored controls are evaluated independently of method ranking. Failed, unfavorable, discordant, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, group-aware partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation. This preserves all values, ordering, dtypes, datasets, endpoints, partitions, methods, predictions, thresholds, statistical procedures, and frozen randomization streams. The correction has no scientific impact on the estimand. All definitive work units are regenerated under the corrected implementation identity; no prior run outputs are reused.
TUNA BİRGÜN· Zenodo (CERN European Organi...· 0 citations
Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It evaluates predictive performance, modality utility, null behavior, failure handling, and compute cost without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. Design and controls All real cohorts were fixed after source and endpoint eligibility checks and before method performance was observed. The design uses shared group-aware partitions, training-only preprocessing, five outer folds repeated three times for real cohorts, 40 independent synthetic replicates, ten target permutations per real cohort, 5,000 paired group bootstraps, and Holm adjustment across the five primary real-dataset contrasts. Literature-anchored controls are evaluated independently of method ranking. Failed, unfavorable, discordant, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, group-aware partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. No deviations are registered at deposit.
TUNA BİRGÜN· Zenodo (CERN European Organi...· 0 citations
Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. The cancer hallmark tool helps in understanding SH3BP2 hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change = 1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-κB and affecting immune responses through WNT/β-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.
This research investigates the application of non-linear constraint optimization to biological systems, specifically focusing on optimization of gene expression and protein folding. Traditional optimization methods often struggle with the inherent complexity and dynamic nature of biological systems, necessitating novel approaches that can effectively capture evolutionary principles and adapt to changing conditions. This paper proposes a hybrid method integrating evolutionary algorithms with constraint optimization, incorporating feedback from the system's own adaptive behavior. We demonstrate the effectiveness of this approach in optimizing a simplified biological model, highlighting its potential for broader applicability in the analysis and control of complex biological processes.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Dynamic Symmetry Field Analysis (DSFA) is a novel technique for unraveling the intricate patterns within complex systems, specifically focusing on identifying and quantifying 'dynamic symmetry fields.' This research explores the potential of employing a new mathematical operator to analyze the evolution of these fields, revealing underlying instabilities and potential for change. The core mechanism leverages the concept of a dynamically evolving symmetry function, allowing for the detection of critical points and deviations from expected behavior. This paper details the mathematical framework, initial results, and potential implications of DSFA for a range of applications, including protein folding, fluid dynamics, and complex mechanical systems. The investigation emphasizes a shift from passive observation to active detection of dynamic structural features, offering a potentially transformative approach to understanding the behavior of such systems.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Here you can find the data from the study "Quadrupling the protein family space with global metagenomics". For Protein Families: iso_clusters25_names.tsv.bz2 Description: TSV file of isolate clusters with more than 25 members. Columns: (1)Cluster name (2)Representative isolate protein header (3) Isolate protein member header (4)Isolate protein member sequence metag_clusters25_names.tsv.bz2 Description: TSV file of metagenomic clusters with more than 25 members. Columns: (1) Cluster name (2) Representative metagenomic protein header (3) Metagenomic protein member header (4) Metagenomic protein member sequence For Protein Folds structures.tar.gz Description: Contains three subfolders with protein structure models HQ:High-Quality (pTM ≥ 0.7) models (26,767 pdb files) MQ:Medium-Quality (0.5 ≤ pTM < 0.7) models(47,823 pdb files) LQ:Low-Quality (pTM < 0.5) models (82,018 pdb files) foldseek_results.tar.gz Description: You will find three folders (HQ, MQ & LQ). Each one contains three files: AF2.tblout (unfiltered hits to AlphaFoldDB) CATH.tblout (unfiltered hits to CATH) PDB.tblout (unfiltered hits to PDB) NMPFAMSDB2_MODELS_SCORES.txt Description: A txt file with pTM and pLDDT score for each family model. Columns: (1) Family name Column (2) pTm score Column (3)pLDDT score
Eleni Aplakidou, Fotis A. Baltoumas, Georgios A. Pavlopoulos· Zenodo (CERN European Organi...· 0 citations
Repeat proteins fold through pathways that are strongly shaped by local energetics, making them highly sensitive to mutations. The ankyrin repeat (AR) domain of IκBα is a cooperative folding unit in which the first four repeats (AR1–AR4) are stable, while the last two are destabilized. Despite this simple modular architecture, IκBα follows a complex folding trajectory involving high-energy intermediates. Here, we investigate how sequence variations modulate folding pathways using coarse-grained AWSEM simulations combined with the Energy Landscape Visualization Method (ELViM) and local frustration analysis. We compare the wild type (WT) with two consensus-designed variants: V93L, which accelerates folding, and L131V, which destabilizes and slows down folding of the protein in vitro. Our results show that WT and V93L share a similar folding funnel, though V93L folds more directly into the native state. In contrast, L131V reshapes the landscape by stabilizing non-native kinetic trap with minimally frustrated contacts. Reversing the L131V mutation allows the protein to reach the native conformation, whereas maintaining it confines the protein to misfolded states that can only be escaped at very high temperatures, without productive folding. These findings highlight how subtle sequence changes can tune frustration, modulate kinetic trapping, and control folding efficiency in repeat proteins.
Murilo N. Sanches, María Inés Freiberger, Peter G. Wolynes et al.· Zenodo (CERN European Organi...· 0 citations
Molecular dynamics (MD) simulations are a fundamental tool in materials science, biology, and pharmaceutical research, offering insights into molecular behavior and dynamics. However, traditional MD simulations often suffer from limitations in accuracy and efficiency, particularly when dealing with complex, dynamic environments. This paper introduces a novel algorithm, termed Adaptive Molecular Dynamics Optimization (AMDO), designed to address these challenges by dynamically adjusting simulation parameters based on a learned model. AMDO leverages an adaptive learning algorithm to optimize the simulation process, resulting in enhanced accuracy and reduced computational cost. We demonstrate the effectiveness of AMDO through the simulation of a complex protein folding process, showcasing improved convergence and reduced simulation time compared to conventional MD methods. The core mechanism centers on the continuous adaptation of simulation parameters, enabling the simulation to effectively capture the nuances of molecular interactions.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Abstract Optical tweezers using metal nanostructures have leveraged extreme subwavelength focusing to circumvent the diffraction limit, enabling the isolation and label-free sensing of single nanoparticles. Over the past two decades, nanoaperture optical tweezers (NOTs) have matured into a useful tool for biophysical analysis, monitoring the conformational dynamics, binding affinities, and structural mutations of single proteins without perturbing labels or tethers. Driven by a post-machine-learning shift in biophysics toward understanding the sequence-structure-dynamics-function paradigm, NOTs have recently achieved real-time mapping of single-protein energy landscapes. Future improvements will aim to resolve the sub-microsecond protein folding dynamics by improving the signal-to-noise ratio while navigating the impacts of surface interactions and thermophoresis. This work forecasts key technical innovations over the next five years to achieve nanosecond-scale temporal resolution. By transitioning to smaller metal nanostructures (which includes moving away from nanoapertures) and operating at longer near-infrared wavelengths, near-field sensitivity can be maximized while mitigating laser-induced heating. Augmented functionalities—including integrated Raman spectroscopy (for applications like peptide identification and single-cell proteomics), and enantioselective chiral trapping will expand the utility of metal nanostructure optical tweezers. Combined with machine learning models to maximize data extraction from low signal-to-noise environments and train future predictive models on protein dynamics, these advances aim to deliver a robust platform to understand the dynamics of biomolecules.
Reuven Gordon· Journal of Physics Photonics· 0 citations
Lignocellulosic biomass is one of the most abundantly available renewable energy sources, but it is entirely underutilized. Composed of a matrix of cellulose and hemicellulose polysaccharides, biochemical deconstruction of this substrate can yield fermentable sugars for biofuel production. Within a typical biorefinery, enzymatic hydrolysis of these polysaccharides is relied upon by Carbohydrate Active enZymes (CAZymes) to cleave the glycosidic linkages between sugar moieties. This technology is significantly hindered though due to the inherent recalcitrance of biomass to enzymatic degradation. This recalcitrance is caused by challenges related to the structural phenolic polymer lignin which limits enzyme accessibility and non-productively binds CAZymes, as well as low overall activity of the soluble enzymes on a highly insoluble, crystalline substrate. Limited accessibility to polysaccharides is typically alleviated by thermochemical pretreatment of biomass, but this does little to overcome the challenges related to CAZyme functionality. Thus, CAZymes and their appended carbohydrate binding modules (CBMs) must be engineered for improved hydrolytic activity. Lignin and cellulose are both found to have slight negative charges on their surfaces, and exploiting these electrostatics using enzyme supercharging is one promising route for improving enzyme activity. Enzyme supercharging refers to the process of mutating several surface exposed amino acid residues to either positive (R, K) or negative (D, E) charged residues to produce high theoretical net charges. Supercharging was applied to both glycosyl hydrolase (GH) CAZymes and their appended CBMs to tune enzyme surfaces to a critical net charge where binding affinity, and ultimately activity, may be optimized. This principle is closely related to the Sabatier principle which plainly states that enzyme turnover is maximized at intermediate strength binding affinity where neither adsorption nor desorption are limiting factors. This design principle was first applied to a GH family 5 endocellulase Cel5A and its native fused family-2a CBM to generate and systematically screen a supercharged library of CBM2a-Cel5A fusions. As a result of supercharging, several supercharged mutants were identified that showed enhanced binding and catalysis with greater than 2-fold improvements in hydrolysis yields. Enzymes with optimized binding affinity additionally exhibit robust thermotolerance in the presence of substrate compared to the native enzyme. Ultimately, strong substrate dependent correlations between net charge and activity were identified, resembling charge related Sabatier optima. Further probes into these effects were performed by applying this supercharging rationale to a GH family 6 exocellulase Cel6B and its native family-2a CBM. Systematic screening identified key differences resulting from supercharging compared to the endocellulase system. Most strikingly, similar Sabatier effects were not identified, yet a key charge engineered enzyme was isolated from the library. Instead of favorable adsorptive properties, this enzyme exhibited altered unfolding pathways between both domains, revealing a 10 °C increase in melt temperature for the catalytic unit. These processive enzymes possess a more complex mechanism for cellulose degradation, and as such, simple affinity modulation was not an adequate predictor of enzyme performance. Exo-endo synergism is a hallmark of efficient biomass utilization, where endo and exo active cellulases work in conjunction to effectively process cellulose. As such, supercharged enzymes of both classes were combined in several ratios to identify how net charge impacts synergistic hydrolysis. In these cases, the enzymes with the highest individual activity were also the best synergistic partners. These findings suggested that electrostatic engineering may provide a broader strategy for modulating enzyme behavior at heterogeneous polymer interfaces beyond lignocellulosic systems alone. To investigate the transferability of these interfacial engineering principles, charge engineered CBMs were applied to poly(ethylene terephthalate) (PET) synthetic polymer hydrolase systems. Several charge engineered CBMs were fused to a Cutinase enzyme that is capable of hydrolyzing ester linkages in the PET backbone. Results found that, while positively supercharged CBMs enhanced PET binding interactions, hydrolysis improvements did not correlate directly with substrate adsorption. Instead, a slightly negative CBM fusion produced substantial improvements in PET depolymerization through enhanced thermostability and catalytic persistence, revealing that enzyme stability, rather than substrate binding, was the dominant limitation governing hydrolysis performance in this system. The synthesis of this work provides a broader framework to the application of protein supercharging as a rational design technique for increasing enzyme activity. Collectively, the findings in this dissertation demonstrate that heterogeneous biocatalysis cannot be reduced to substrate binding interactions alone. Instead, enzyme activity improvement depends on a dynamic balance between substrate engagement, catalytic turnover, enzyme stability, and catalytic persistence at insoluble polymer interfaces. These results establish electrostatic surface engineering as a powerful, but highly context-dependent strategy for modulating enzyme behavior across diverse heterogeneous polymer degradation systems.
Antonio De Chellis· Rutgers University Community...· 0 citations
Non-additive interactions between environmental stressors, where biological responses to simultaneous stressors do not equal the sum of individual-stressor responses, commonly occur across organisms and environments. To enable predictions of organismal resilience to shifting patterns of environmental stressors, it is important to both identify these interactions and document the mechanisms underlying them. The supratidal copepod Tigriopus californicus demonstrates a non-additive, antagonistic pattern of increased heat tolerance when simultaneously exposed to high salinities. We investigated salinity’s chronic and acute effects on heat tolerance in a northern population and quantified responses in protein abundance to hypersalinity and heat stress, both alone and in combination. The overall proteomic response to multiple stressors was non-additive and largely reflected that to high temperature. However, 42% of multi-stressor proteins were absent from either single-stressor response; we refer to these proteins that are only differentially abundant in the multi-stressor scenario as “emergent”. Our results suggest that the increased heat tolerance of T. californicus conferred by hypersalinity may be driven by a combination of these emergent proteins, several proteins induced by hypersalinity in both single- and multi-stressor conditions that may contribute to cross-tolerance, and four proteins with additive abundance patterns (including a small heat shock protein). These candidate proteins play putative roles in several relevant processes including the heat shock response, protein folding, and regulation of metabolism and oxidative stress responses. Our results connect to prior whole-organism findings and highlight promising pathways for future investigation in the context of heat tolerance and multi-stressor interactions.
C. Terry, Maxime Leprêtre, Dietmar Kültz et al.· Physiological Genomics· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.