Skip to content
Open access

Machine Learning Gap-Fills Missing Transporter Kinetics in Biosystems Across Scales

Jul 2026 · bioRxiv · 0 citations
Biology

TL;DR

The first compound-protein interaction machine learning model of transporter Vmax and Km, MMTKPred is developed, which captures the effects of point mutations and substrate changes on transporters and effectively models metabolite transport spanning from molecular to multi- species scales.

Abstract

Understanding transporter kinetics is essential for deciphering metabolite exchanges in biosystems, particularly for cells subject to substrate gradients. Nevertheless, the prediction of transporter kinetic parameters, maximum rate per gram protein (Vmax) and Michaelis-Menten constant (Km), has not yet been tackled. Here, we developed the first compound-protein interaction machine learning model of transporter Vmax and Km, MMTKPred, which achieved R2=0.553, RMSE=1.155 mmol/hr/g Protein and R2=0.330, RMSE=0.935 mM for log10-scaled Vmax and Km prediction, respectively. Moreover, we demonstrated MMTKPred’s predictive power across biosystem scales, from capturing transporter kinetics modulated by point mutations and substrate changes at the molecular level, to enabling substrate-sensitive metabolic modelling of non-model yeasts at the cellular level, and rationalizing inter-species substrate competition in co-cultures. Collectively, MMTKPred effectively models metabolite transport spanning from molecular to multi- species scales, thereby offering a computational tool for rational microbial cell factory optimization. Graphical abstract Highlights MMTKPred, first transporter kinetics CPI model, reaches ∼1 log10 RMSE for Vmax and Km. MMTKPred captures the effects of point mutations and substrate changes on transporters. Predicted kinetics enables substrate sensitivity in metabolic flux modelling. Predicted kinetics explains inter-species substrate competition outcomes.

Read PDF

Similar papers

Open access Jul 2026

An enzyme-specific protein language model for catalytic property prediction

This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.

Chong Wang, Mengyao Li, Shaolei Geng et al. · 0 citations
Open access Jul 2026

A hybrid machine learning and enzyme-constrained metabolic model for ab initio prediction of proteome reallocation

High expression of heterologous proteins in microbial cell factories frequently triggers a severe burden due to reallocation of finite cellular proteome. Conventional constraint-based models struggle to predict these resource shifts ab initio without relying on condition-specific omics data. To bridge this gap, we developed the Hybrid Transcription-Translation (HyTT) framework, combining multivariate adaptive regression splines (MARS) with enzyme-constrained metabolic models by enforcing an 80S ribosome integrity constraint. Cast as a mixed-integer linear programming problem, HyTT mathematically couples macroscopic spatial boundaries with microscopic, sequence-derived translational costs based on a bisection search. Validation against steady-state chemostat quantitative proteomics data demonstrated the superior capability of HyTT over contenders in predicting system-wide resource (re)allocation in Saccharomyces cerevisiae. Operating ab initio, the framework doubled the predictive accuracy of protein abundances (Pearson r=0.501) compared to conventional models, successfully segregating the minimal essential proteome from the cellular reserve pool. Crucially, HyTT autonomously captures complex stress responses vital for metabolic engineering. Upon simulating a 15% recombinant protein burden, the framework accurately predicted systemic growth retardation, decrease of ribosomal portion of the proteome, and surge of ethanol production, in line with the Crabtree effect. System-level analysis uncovered that cells adapt to restricted proteomic capacity through non-uniform metabolic rerouting, downregulating respiratory complexes in favor of high-turnover glycolytic enzymes, and relying on ribosomal paralog switching to minimize sequence-specific assembly costs. Ultimately, HyTT provides a computationally agile, sequence-driven platform for decoding dynamic resource reallocation, offering a powerful predictive tool to navigate metabolic trade-offs and guide rational strain design without requiring condition-specific multi-omics inputs.

E. Motamedian, Z. Nikoloski · 0 citations
Open access Jul 2026

Machine learning guided cell-free expression maps the biochemical landscape of carbonic anhydrase

This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.

J. Lazar, Evan Komp, I. Martínez et al. · 1 citation
Open access Jul 2026

Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

A machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round that may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants.

Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al. · 0 citations
Open access Jul 2026

WILDkCAT: extract, retrieve, and predict enzyme turnover numbers of constraint-based metabolic models

Abstract Summary Accurate enzyme turnover numbers are essential for building enzyme-constrained genome-scale metabolic models. However, collecting and curating these parameters remains a major bottleneck. Indeed, kcat values are scattered across multiple databases, reported under varying experimental conditions, and often missing for many enzymes. To address this challenge, we present WILDkCAT, a Python-based pipeline that enables the retrieval of kcat values from wild-type enzyme measured under user-specified pH and temperature ranges for a given metabolic model. The application to Escherichia coli (iML1515) and Homo sapiens (Human-GEM) models demonstrated the ability of WILDkCAT to retrieve substantial kcat coverage and its applicability across diverse genome-scale models. Availability and implementation WILDkCAT is available at https://github.com/sysbiolux/WILDkCAT and from PyPI. WILDkCAT works on all major operating systems and computer architectures. The documentation is available at https://sysbiolux.github.io/WILDkCAT.

Hugues Escoffier, Carole L. Linster, Thomas Sauter · 0 citations