The accuracy of the computational estimation of relative free energies (e.g., for solvation or protein–ligand binding) depends on the smoothness of the phase-space transformation between the two alchemical end-states. A smooth transformation ensures sufficient phase-space overlap between the neighboring intermediate states connecting the two end-states in equilibrium (EQ) simulations and generates less dissipative work in nonequilibrium (NEQ) simulations. The conventional energy interpolation (EI) coupling scheme constructs the intermediate states by linearly combining the end-state potentials. We show that the enveloping distribution sampling (EDS) coupling scheme, a generalization of EI where the corresponding Boltzmann factors are linearly combined, represents a much more flexible alternative. Through the use of a negative smoothing parameter, the EDS scheme increases the local curvature of the sampling phase space along the transformation axis, thereby avoiding phase transitions and creating a smoother transformation. We validate this behavior in increasingly complex settings, from harmonic oscillators and Ising model systems to absolute hydration free-energy (AHFE) calculations on the FreeSolv data set. EDS consistently yields more accurate and statistically robust free-energy estimates compared to the conventional EI scheme for the model system calculations, while a clear advantage is observed for AHFE in the NEQ regime, where less dissipative transitions lead to more reliable free-energy estimates.
Shu-Yu Chen, Enrico Ruijsenaars, P. Hünenberger et al.· Journal of Chemical Theory a...· 0 citations
Neural network potentials (NNPs) can provide insight into biological processes at atomic resolution. Training these NNPs requires large and diverse datasets of molecules, conformations, and configurations. However, so far little attention has been paid to the description of solvation, despite its importance for biomolecular systems. This work lays the foundation for NNPs where solvation is an integral part of the model. Following a quantum-mechanics/molecular-mechanics (QM/MM) formalism with an electrostatic embedding scheme, systems are decomposed into a QM zone with the solute(s), which is electrostatically coupled to the point charges from surrounding solvent molecules (MM zone). Using an accelerated sampling approach, we generate the biomolecular multiscale simulation (BMS25) dataset with over 50,000 topologies and more than 1.5 million unique conformations of peptides and miniproteins as well as small molecules and transition states from chemical reactions. The dataset includes energies, gradients, and multipoles of solute molecules as well as gradients on solvent molecules at the
ω
B97M-D4/ma-def2-TZVPP level of theory, enabling the development of multiscale NNPs for simulating large biomolecular systems.
Moritz Thürlemann, Felix Pultar, Igor Gordiy et al.· Scientific Data· 0 citations
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
Idil Ismail, Gregory A Landrum, Sereina Riniker· Journal of the American Chem...· 0 citations