Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

ProtSyntax: a protein large language model for decoding post-translational modification syntax and function

Post-translational modifications (PTMs) expand protein function by encoding context-dependent regulatory states, and their dysregulation contributes to cancer, neurodegeneration and metabolic disease. However, existing methods treat PTMs as independent residue labels, limiting their ability to distinguish contextually permissible sites, model crosstalk and infer functional consequences. Here we introduce ProtSyntax, a PTM-aware foundation protein language model combining protein-aware positional encoding, bidirectional state-space propagation, geometry-constrained attention and adaptive multi-objective learning. This design integrates residue chemistry, motif order, long-range context and three-dimensional microenvironments while coupling PTM recognition to enzyme function. Across 40 PTM-site benchmarks, ProtSyntax exceeded the strongest baselines in mean MCC and AP by 12.66% and 10.67%. ProtSyntax also recovered masked PTM types and sites, rejected structural decoys, generalized to data-scarce modifications, reconstructed crosstalk and linked PTM perturbations to enzyme kinetics. Applications to pathogenic variants, biomolecular condensates and disease-associated PTM landscapes demonstrate its potential to decode the regulatory language of the modified proteome.

Yiyu Lin, Jiahui Wu, You Zhou et al. · 0 citations
Open access Jul 2026

Fine-tuned large language models enhance influenza forecasting.

Influenza-like illness (ILI) remains a persistent global health challenge, necessitating accurate forecasting tools for timely public health response. This study systematically benchmarks fine-tuned large language models (LLMs), e.g., Llama2 and GPT2, for influenza surveillance forecasting in data-limited time-series settings. We develop a lightweight fine-tuning framework that adapts pre-trained LLMs using compact embedding and prediction layers and evaluate it on seven weekly aggregated real-world surveillance datasets. Despite sample sizes of only ∼523 time points per region and the absence of cloud-based data transfer, fine-tuned LLMs consistently outperform SARIMA, LSTM, PatchTST, CoVTransformer, FEDformer, Time-LLM, and GPT4TS in both accuracy and stability, especially for long-term forecasts across diverse geographic settings. Even in zero-shot settings, pre-trained LLMs capture broad epidemic trends with performance comparable to SARIMA. These findings establish fine-tuned LLMs as efficient and robust forecasting tools suitable for privacy-sensitive, data-scarce public health applications.

Chenxi Li, Wenjing Gao, Qiqiao Zhang et al. · 0 citations