Glucoamylases are highly demanded by the industry for starch saccharification, which necessitates an increase in enzyme activity. To improve the activity of Aspergillus awamori VKPM F-1262 glucoamylase N181Q (AGAM), amino acid substitutions were made in the region of substrate binding. In the case of the variant D237G, there was a 1.3–1.4-fold increase in activity towards soluble starch for the individual enzyme and crude enzyme solution after cultivation in flasks. For the variants G57A and C319T, there was also a 1.1–1.4-fold increase in enzyme activity for the crude flask solution, while the increase in individual enzyme activity was minor. AGAM showed an increase in apparent molecular weight consistent with dimerization in the presence of soluble starch. A high T optimum of 70 °C was observed for AGAM and the variants in the hydrolysis of soluble starch. In further protein engineering, a high T optimum for hydrolysis, which may coincide with dimerization, may be considered as a hypothesis for the selection of thermostable glucoamylases, while an area of binding of oligo- and polysaccharides may be a site for mutagenesis.
Anna S. Dotsenko, Nikita Eroshenko, Ekaterina Rubtsova et al.· BioTech· 0 citations
Eukaryotic translation initiation factor 5A (eIF5A) is a highly conserved protein family unique to eukaryotes, yet its functional characterization in woody plants remains limited. In this study, we identified four eIF5A genes (PtoeIF5A1–PtoeIF5A4) from the genome of Populus tomentosa, a fast-growing tree species indigenous to China, and characterized their expression patterns and functional roles through bioinformatics analysis, quantitative real-time PCR, stable overexpression in Arabidopsis thaliana, and transient expression in Nicotiana benthamiana leaves. Our results demonstrated that all PtoeIF5A proteins contain a conserved OB-fold domain and multiple phosphorylation sites, with PtoeIF5A1 showing predominant expression in roots and secondary xylem. Functional assays revealed that PtoeIF5A1 overexpression accelerated inflorescence stem elongation and early flowering in Arabidopsis, induced visible chlorosis and programmed cell death (PCD) in tobacco leaves, and significantly enhanced salt tolerance under NaCl treatment. Collectively, these findings establish PtoeIF5A1 in poplar as a pleiotropic regulator integrating developmental cues, programmed cell death, and stress responses; and as a valuable genetic resource for breeding stress-resilient woody plants.
Dan Zhu, Guan Yang, Feng Feng et al.· Plants· 0 citations
Abstract Probiotic efficacy depends not only on gastrointestinal survival but on mucosal adhesion and the capacity to deliver bioactive molecules at the intestinal surface. This study optimized a whey protein isolate (WPI)–chitosan (CS) matrix for spray drying microencapsulation of Lacticaseibacillus casei BL23 using central composite design, targeting enhanced mucoadhesion while preserving bacterial viability and extracellular vesicle (EV) secretion capacity. The optimal formulation, WPI 20%-CS 0.5%, yielded viable counts within recommended probiotic ranges (6.6 × 10⁹ CFU/g) and a ~ 77-fold mucoadhesion increase relative to WPI alone, supporting extended intestinal residence time and potentially enhanced therapeutic efficacy. Storage stability was confirmed at 4°C and − 20°C, and microcapsules were predominantly spherical (2–15 μm), suitable for food applications. Microencapsulation also significantly enhanced gastrointestinal survival, with encapsulated bacteria showing only a ~ 2-log reduction after the gastric phase compared to ~ 6-log for free bacteria. Fermentative capacity in reconstituted milk was fully preserved, with reduced syneresis indicative of improved gel stability. Critically, this system ensures the delivery of probiotic-derived extracellular vesicles (EVs) at the intestinal interface. We demonstrate that EV secretion is maintained post-encapsulation, yielding vesicles (70–100 nm) enriched in p40 and p75 proteins. To our knowledge, this is the first evidence that spray-dried microcapsules can effectively serve as a delivery platform for postbiotic EVs by preserving the functional secretory machinery of the encapsulated bacteria. These findings position mucoadhesive WPI-CS microcapsules as a robust strategy for the targeted delivery of EVs in nutraceutical and functional food development. Key points • Whey protein-chitosan microencapsulation boosts mucoadhesion ~77-fold in vitro. • Microencapsulation enhances probiotic survival across gastrointestinal conditions. • Spray-drying preserves the probiotic machinery for postbiotic EV secretion.
Cecilia L. D’Antoni, Rocío Corfield, Sergio I. Nemirovsky et al.· Applied Microbiology and Bio...· 0 citations
AIM: Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. MATERIALS AND METHODS: The cancer hallmark tool helps in understanding SH3BP2 hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. RESULTS AND CONCLUSIONS: The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change = 1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-κB and affecting immune responses through WNT/β-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.
Protein language models (pLMs) learn sequence patterns at evolutionary scale, but these patterns remain inaccessible within these “black box” models. To discover them, we developed MotifAE, an unsupervised framework based on the sparse autoencoder (SAE) architecture that projects pLM embeddings into an interpretable, sparse latent space. MotifAE introduces an additional smoothness loss to encourage coherent feature activation, which markedly improves the identification of known functional motifs compared to the standard SAE. The sequence patterns captured by MotifAE exhibit rich diversity, align with known functional motifs, and are reflected in the model’s weight space. Beyond short motifs, MotifAE also captures some structural domains, with latent feature activation scores correlating with residue importance for diverse domain functions. By aligning MotifAE features with experimental data, we further identified features associated with domain folding stability. These features enable the prediction of a stability-specific fitness landscape. Overall, MotifAE provides a general framework for systematic sequence pattern discovery and interpretation, with the potential to advance protein function analysis, mutation effect interpretation, and rational protein engineering. Protein language models (pLMs) capture biologically meaningful sequence patterns, but the features learned by these models remain largely opaque. Here the authors present MotifAE, an interpretable sparse autoencoder framework that reveals diverse sequence motifs and structural domains from pLM embeddings, improves recovery of known functional motifs and identifies features linked to domain function and folding stability, enabling stability-specific fitness landscape prediction.
Chao Hou, Ди Лю, Yufeng Shen· Nature Communications· 0 citations
This is an updated version of a presentation at Edinburgh University in June with some editing and additional slides on BindingDB patent curation Presented at the Scaggs School of Pharmacy, University of California San Diego, August 28th 2026 as guest of Prof. Michael Gilson Abstract The beta amyloid (APP) cleaving enzyme (BACE1) was identified as a drug target for Alzheimer\'s Disease (AD) in 1999 while its paralog (BACE2) was proposed as a target for type II diabetes (T2DM) in 2011. The generation of BACE1 inhibitors over ~ 25 years has made it one of the most intensely persued AD drug targets, with lead compounds curated by the Guide to Pharmacoly, liteature inhibitors by ChEMBL and compounds from patents by BindingDB. Unfortunately, no less than six small-molecule BACE1 inhibitors have failed in Phase II or III clinical trials for AD. Despite reducing Aβ production, several programs reported cognitive and neuropsychiatric problems, raising questions on the “normal” roles of BACE1. To shed some light on these an evolutionary analysis was undertaken in 2013 (PMID: 24381583). This identified single-copy homologs (UrBACE) wirh 35-45% protein sequence identity to mammalian BACE1 in many basal animal phyla. More homologues have recently been identified from new molluscan genomes, thereby extending support for an evolutionary histroy as a duplication of the UrBACE in fish giving rise to BACE1 and BACE2 paralogues in all vertebrate lineages. In addition, Alpha fold structures for the Sea Urchin UrBACE have been generated. However, it remains unclear what functional roles this enzyme had both before and during evolution of the ancestral nervous system over ~ 800 million years ago. By illuminating UrBACE roles, functional genomics experiments may shed on the clinical failure of human BACE1 inhibitors in AD.
Southan Christopher· Zenodo (CERN European Organi...· 0 citations
Abstract Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular “hijacking”, Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana , contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a swapped domain topology; the other has no disulfide bonds and possesses two distinct, tandem domains. To explore the structural evolution of bicycle proteins, we attempted to predict bicycle protein structures with Alphafold2 (AF2) and other deep learning programs. While AF2 did not recover the two experimental structures using existing databases, it succeeded when provided with multiple sequence alignments (MSAs) of protein sequences from newly sequenced closely related species. Using this approach, we generated 2400 high-confidence bicycle protein predictions from seven aphid species. While all aphid bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance. Our extension of AF2 with custom MSAs of proteins from closely related species provides a generalizable, powerful approach for predicting structures of rapidly evolving protein families. Significance statement Parasites introduce specialized “effector” proteins into hosts to suppress host immunity and to release nutrients. The molecular functions and structures of most effector proteins are unknown. Effector proteins often evolve rapidly and share no similarity with proteins of known function. Here, we demonstrate that machine learning algorithms can predict the structures of aphid “bicycle” effector proteins when supplemented with data from closely related species. We exploit this finding to generate predictions of 2400 bicycle protein structures. Aphid bicycle proteins exploit a common folding motif, yet exhibit topologically distinct structures that form separate structural clusters. Despite the clustering of these proteins in structure space, they occupy a nearly uniformly physicochemical space, suggesting that they encode a large diversity of molecular functions.
Fatema Bhinderwala, Aishwarya Korgaonkar, Kota N. Gopalakrishna et al.· Proceedings of the National...· 0 citations
Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It tests predictive performance, modality utility, model-capacity effects, missingness handling, null behavior, exploratory external transport, and failure handling without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families with operative structured missingness for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. CPTAC UCEC exploratory holdout: transcriptomics and copy number for endometrioid-versus-serous transport assessment across 95 patient-disjoint tumors. Methods and controls Nine fixed methods compare Omicau with an unmasked architecture-matched ablation, matched early and single-modality neural controls, early and single-modality linear controls, weighted late fusion, a missingness-only diagnostic, and calibrated latent partial least squares. A TCGA-UCEC complete-training-feature sensitivity tests outcome-associated technical missingness. Internal cohorts use shared group-aware partitions, training-only preprocessing, five outer folds repeated three times, 5,000 paired group bootstraps, paired DeLong tests for AUROC, 4,999 paired squared-error sign flips for R-squared, and Holm adjustment across five primary matched-capacity contrasts. Effect sizes and intervals are the primary evidence. Because repeated out-of-fold predictions share training sets, internal p-values are conditional on the frozen prediction vectors and are not unconditional population-generalization tests. Ten target permutations per cohort are coarse catastrophic-leakage diagnostics, not formal empirical tail-probability tests. TCGA and CPTAC expression scales are harmonized by a source-declared, target-blind transformation with pooled and matched-feature numerical-domain gates. The CPTAC endpoint is exploratory because it was exercised during predeposit development smoke. Literature-anchored controls remain independent of method ranking. Failed, unfavorable, discordant, non-estimable, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, internal and external partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation while preserving scientific values and frozen randomization streams. Deviation 2 - Comparator and ablation expansion. Matched neural, unmasked, missingness-only, complete-feature, weighted late-fusion, and partial least-squares controls separate fusion value from model capacity, technical missingness, and integration strategy. All settings are fixed before definitive execution. Deviation 3 - Independent external evaluation. A CPTAC UCEC holdout adds 95 patient-disjoint assessment cases. TCGA-UCEC supplies all training and model-selection rows; shared transcriptomic and copy-number features are aligned by unique Entrez identifiers and expression scales are harmonized without using CPTAC outcomes. Deviation 4 - Primary contrast realignment. The primary contrast is Omicau minus the matched early neural control. Prior linear comparisons remain fully reported as contextual estimates and are not substituted for the matched-capacity test. Deviation 5 - External development exposure. The CPTAC endpoint was exercised during predeposit development smoke. External estimates are designated exploratory and are not treated as untouched confirmatory validation. Deviation 6 - External expression-scale correction. Predeposit development smoke exposed incompatible TCGA and CPTAC expression domains. Linear TCGA RSEM values now receive log2(x+1) after negative values are marked missing; CPTAC retains its source-declared log2 scale. Target-blind pooled and matched-feature gates validate compatibility. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 7 - Synthetic missingness application correction. Predeposit audit showed that registered synthetic missingness masks were not reaching model matrices. The masks now alter every synthetic method input exactly as registered. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 8 - Numerically constant diagnostic correction. The published 2.0.0 external missingness-only control had identical inputs but machine-precision score differences that produced spurious rank metrics. Version 2.0.1 assigns identical assessment rows an exact shared score and treats numerically constant classification scores as non-discriminating. All definitive results are regenerated under the corrected identities. The primary Omicau contrasts are unchanged by this correction.
TUNA BİRGÜN· Zenodo (CERN European Organi...· 0 citations
Overview A prospectively frozen benchmark suite for leakage-safe multi-omic integration. It tests predictive performance, modality utility, model-capacity effects, missingness handling, null behavior, exploratory external transport, and failure handling without making claims of clinical utility or causal biological inference. Included datasets Synthetic controls: paired null and planted-signal families with operative structured missingness for binary classification and continuous regression. DepMap/CCLE: transcriptomics, copy number, LC-MS metabolomics, and PRMT5 dependency across 644 cell lines. TCGA BRCA: transcriptomics, copy number, and RPPA protein abundance for ductal-versus-lobular classification across 783 tumors. TCGA LGG, KIRC, and UCEC: transcriptomics and copy number for IDH status, pathological stage, and histology endpoints across 507, 507, and 500 tumors, respectively. CPTAC UCEC exploratory holdout: transcriptomics and copy number for endometrioid-versus-serous transport assessment across 95 patient-disjoint tumors. Methods and controls Nine fixed methods compare Omicau with an unmasked architecture-matched ablation, matched early and single-modality neural controls, early and single-modality linear controls, weighted late fusion, a missingness-only diagnostic, and calibrated latent partial least squares. A TCGA-UCEC complete-training-feature sensitivity tests outcome-associated technical missingness. Internal cohorts use shared group-aware partitions, training-only preprocessing, five outer folds repeated three times, 5,000 paired group bootstraps, paired DeLong tests for AUROC, 4,999 paired squared-error sign flips for R-squared, and Holm adjustment across five primary matched-capacity contrasts. Effect sizes and intervals are the primary evidence. Because repeated out-of-fold predictions share training sets, internal p-values are conditional on the frozen prediction vectors and are not unconditional population-generalization tests. Ten target permutations per cohort are coarse catastrophic-leakage diagnostics, not formal empirical tail-probability tests. TCGA and CPTAC expression scales are harmonized by a source-declared, target-blind transformation with pooled and matched-feature numerical-domain gates. The CPTAC endpoint is exploratory because it was exercised during predeposit development smoke. Literature-anchored controls remain independent of method ranking. Failed, unfavorable, discordant, non-estimable, and indeterminate outcomes remain reportable. Reproducibility The archive contains the frozen protocol, immutable source registry, download and validation code, internal and external partitions, fixed comparator implementations, statistical aggregation, schemas, environment pins, and fault-injection tests. Raw molecular matrices, participant-level data, local paths, and benchmark results are excluded. Deviations Deviation 1 - Aggregation target normalization. Final aggregation converts read-only NumPy memory-mapped target vectors to base NumPy arrays before metric and bootstrap validation while preserving scientific values and frozen randomization streams. Deviation 2 - Comparator and ablation expansion. Matched neural, unmasked, missingness-only, complete-feature, weighted late-fusion, and partial least-squares controls separate fusion value from model capacity, technical missingness, and integration strategy. All settings are fixed before definitive execution. Deviation 3 - Independent external evaluation. A CPTAC UCEC holdout adds 95 patient-disjoint assessment cases. TCGA-UCEC supplies all training and model-selection rows; shared transcriptomic and copy-number features are aligned by unique Entrez identifiers and expression scales are harmonized without using CPTAC outcomes. Deviation 4 - Primary contrast realignment. The primary contrast is Omicau minus the matched early neural control. Prior linear comparisons remain fully reported as contextual estimates and are not substituted for the matched-capacity test. Deviation 5 - External development exposure. The CPTAC endpoint was exercised during predeposit development smoke. External estimates are designated exploratory and are not treated as untouched confirmatory validation. Deviation 6 - External expression-scale correction. Predeposit development smoke exposed incompatible TCGA and CPTAC expression domains. Linear TCGA RSEM values now receive log2(x+1) after negative values are marked missing; CPTAC retains its source-declared log2 scale. Target-blind pooled and matched-feature gates validate compatibility. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 7 - Synthetic missingness application correction. Predeposit audit showed that registered synthetic missingness masks were not reaching model matrices. The masks now alter every synthetic method input exactly as registered. The correction precedes definitive execution, and earlier smoke outputs are not reused. Deviation 8 - Numerically constant diagnostic correction. The published 2.0.0 external missingness-only control had identical inputs but machine-precision score differences that produced spurious rank metrics. Version 2.0.1 assigns identical assessment rows an exact shared score and treats numerically constant classification scores as non-discriminating. All definitive results are regenerated under the corrected identities. The primary Omicau contrasts are unchanged by this correction.
TUNA BİRGÜN· Zenodo (CERN European Organi...· 0 citations
Across every scale at which we study the brain, from folded proteins and single neurons to cortical populations and the moving body, artificial intelligence (AI) has shifted from a bespoke tool into a driver of measurement and, increasingly, a generative engine for hypotheses. Here, I review recent progress (and open challenges) in applying AI for neuroscience along this scale axis: structure prediction for proteins, simulation-based inference for biophysical neurons, latent and dynamical models for neural populations, task-trained networks as minimal models of circuit computation, computer vision for animal behavior, neuromusculoskeletal modeling for biomechanics, and the multimodal, agentic systems now promising to automate discovery itself.
Mackenzie Weygandt Mathis· Current Opinion in Neurobiol...· 0 citations
Coastal waters off Nouakchott support fisheries and, more recently, offshore oil activity and iron-ore export, all of which raise the risk of marine contamination. Local biomonitoring tools are still lacking. We asked whether the wedge clam Donax rugosus , abundant on Nouakchott’s beaches, can serve as a sentinel. Acetylcholinesterase (AChE), an enzyme inhibited by organophosphate and carbamate pesticides and by several trace metals, was measured together with the physiological condition index (CI). Clams were collected monthly from September 2022 to July 2023 at two beaches, Nicola and Hôtel Sabah. AChE was first assayed in four organs (mantle, digestive gland, foot, gill) and then followed monthly in the gills. Activity differed strongly among organs: the mantle was highest (about 16.0 nmol min⁻¹ mg⁻¹ protein, two-site mean) and the gill lowest (about 2.6), roughly a six-fold range. We retained the gill for the time series because its low, stable baseline is easy to read. Gill AChE and CI both changed with season, each peaking in summer; gill AChE was lowest in autumn and rose with water temperature, while CI fell in spring, the period when Donax spawns. Paired tests showed no significant difference between the two beaches for either variable, and the AChE–CI link was weak overall. Gill AChE in D. rugosus is a workable, low-baseline biomarker for coastal biomonitoring in Mauritania, on the condition that temperature and reproductive state are accounted for when the data are interpreted.
Roughaya Hassen Sarr, Sidi Ahmed Elemin, Z. Sidoumou et al.· Journal of Ecological Engine...· 0 citations
Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.
Yifan Feng, Guang Cheng, Shihui Ying et al.· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.