Skip to content
Open access

WaterFlow: Prediction of Ordered Water Molecule Positions on Protein Structures

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

WaterFlow is introduced, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures that outperforms the existing state of the art at every precision level and quantifies the tradeoff between data quantity and data quality.

Abstract

Ordered water molecules mediate many protein functions including stability, ligand binding, and catalysis. Predicting their positions with sub-angstrom accuracy would support protein design, binding affinity prediction, and automated model building in X-ray crystallography and cryo-EM. However, water molecule prediction lags behind protein and other molecule structure predictions. We introduce WaterFlow, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures. WaterFlow outperforms the existing state of the art at every precision level. We demonstrate that this model not only predicts ground truth modeled water molecules, including those around ligands, but also fits the underlying experimental data well, and therefore proposes that it may be used for both prediction and modeling water molecules. We demonstrate that WaterFlow’s novel predictions are often associated with positive electron difference density, meaning the model places water molecules at sites the original structure depositions omitted. We use this improved model to address the data constraint. By mapping the Pareto front of achievable accuracy of water molecule prediction, alongside analysis of different training data schemas, we quantified the tradeoff between data quantity and data quality, demonstrating the diversity of high quality structures is limiting the results possible. Overall, WaterFlow predicts ordered water to serve as a solvent module for structure-based drug design, and predicted structures, as well as for water molecule placement during crystallographic refinement.

Read PDF

Similar papers

Does Surface Conservation Yield? Application to Data-Driven Docking

The interface prediction program WHISCY is presented, which combines surface conservation and structural information to predict protein–protein interfaces and demonstrates the potential of using interface predictions to drive protein–protein docking.

Sjoerd J. de Vries, A. V. van Dijk, A. M. Bonvin · 0 citations
Open access Aug 2026

PreFold-dG: estimating binding affinity of protein–protein interaction from intermediate representations of protein folding model

Abstract Motivation Binding affinity governs how proteins interact and underlies essential biological processes. Computational approaches have been developed to simulate and predict protein binding, but the scarcity of high-quality data has imposed significant constraints. One consequence is that most methods focus on predicting mutational changes in binding affinity (ΔΔG), rather than binding affinity (ΔG) itself. This practice risks overfitting to skewed data distributions, limiting the generalizability of predictions. Recent advances in protein structure prediction have enabled computational modeling of protein conformations in mass, providing rich structural information from which binding interactions can be largely explained. However, leveraging these advances for effective prediction of binding affinity has yet to translate into reliable predictions. Results We present PreFold-dG, a model that estimates binding affinities of protein complexes utilizing intermediate embeddings from Boltz-2, an open-source foundation model for protein structure prediction. Our approach aggregates residue-level information weighted by interresidue distance, and predicts ΔG directly rather than its derivative, ΔΔG. PreFold-dG achieved state-of-the-art performance on well-established binding affinity prediction benchmarks and demonstrated robustness on independent test sets. Ablation studies suggest that all intermediate embeddings are utilized in the prediction, whereas their contributions to modeling ΔΔG and ΔG vary. We further validated our model through case studies on real-world broadly neutralizing antibody data with evolutionary relevance. Availability https://github.com/LGAI-Research/PreFold-dG.

Sungjoon Park, Soorin Yim, Dongyun Kim et al. · 0 citations
Open access Aug 2026

Conserved water molecules shape the pathogenicity of missense variants in human proteins.

Conserved water molecules (CWMs) are tightly bound solvent molecules that occupy well-defined, recurrent positions in protein structures. Although they are known to influence protein stability, function, and ligand binding, their role in shaping the effects of human missense variants remains largely unexplored. Here, we demonstrate that CWMs are a previously underappreciated determinant of missense variant pathogenicity. By predicting ligand-binding and CWM sites across human PDB structures and mapping missense variants to these sites and the remaining protein surface, we found that pathogenic variants were significantly enriched at CWM sites, whether overlapping or outside other ligand-binding regions. This enrichment exceeded that observed for binding sites as a whole, indicating a broader role for water-mediated interactions in modulating variant effects. To explore a mechanistic basis for this association, we performed molecular dynamics simulations of human lysosomal acid glucosylceramidase (GCase), encoded by GBA1 and implicated in Gaucher disease and Parkinson's disease risk. Selective destabilization of a CWM site in wild-type GCase produced structural and dynamical changes resembling those observed in the pathogenic L444P variant, whereas stabilization of this site in L444P shifted several measures toward wild-type behavior. These results suggest that disruption of a single CWM can contribute to long-range structural remodeling observed in a disease-associated variant. Together, our findings identify CWMs as a novel structural constraint shaping the distribution and effects of pathogenic missense variants. Incorporating water-mediated interactions into structural models provides a generalizable framework for interpreting human genetic variation and its contribution to disease.

Janez Konc, Karmen Recer, Tanja Kunej et al. · 0 citations
Review Open access Jul 2026

Solvent Interaction Analysis: A New Lens for Protein Structure and Diagnostics

Aqueous two-phase systems (ATPSs) provide a versatile, fully aqueous platform for probing solute–water interactions and protein structure. This review first surveys the diversity and phase behavior of biphasic aqueous systems formed by polymers and salts. We describe how phase diagrams characterize ATPS formation and composition and how both polymer chemistry and salt identity, rather than molecular size alone, govern phase separation by modulating the solvent properties of water. Building on a modified binodal model, we show that phase separation and solute partitioning can be understood in terms of changes in aqueous solvent dipolarity/polarizability, hydrogen-bond donor/acceptor properties, hydrophobicity, and electrostatics, quantified via solvatochromic probes and homologous solute series. These measurements underpin solvent interaction analysis (SIA), in which the partition coefficients of small molecules and proteins across panels of ATPSs are used to generate “structural signatures” that sensitively report on amino acid substitutions, conformational changes, aggregation, ligand binding, osmolyte effects, and post-translational modifications, independent of protein size. We discuss how SIA can be implemented in vial-, plate-, and microfluidic formats and combined with diverse analytical readouts (HPLC, MS, colorimetric assays, and immunoassays), and we contrast this structure-focused approach with conventional concentration-only proteomic and biomarker strategies. Particular emphasis is placed on structure-based biomarker discovery, where disease-relevant shifts in proteoform distributions—especially glycosylation changes—are often more informative than bulk protein levels and where SIA can complement or simplify complex glycomics and top-down proteomics workflows. As a case study, we describe the recently FDA-approved IsoPSA assay, which applies SIA principles to prostate-specific antigen by measuring cancer-associated structural alterations in circulating PSA via its partition behavior in a proprietary ATPS. IsoPSA generates a single index that discriminates between high-grade prostate cancer and benign and low-grade conditions. Prospective, longitudinal, and MRI-integrated clinical studies demonstrate that IsoPSA improves pre-biopsy risk stratification, reduces unnecessary biopsies, and provides robust negative and positive predictive values within the PSA “gray zone.” Collectively, the data support aqueous solvent interaction analysis as a broadly applicable, mechanistically grounded technology for protein characterization, drug–protein interaction studies, and structure-centric biomarker development, exemplified by the clinical translation of IsoPSA.

B. Zaslavsky, M. Stovsky, V. Uversky · 0 citations