bioRxiv Science⌕ Search

Biology subjects

HOLTON, J. M.

Publications and source records attributed to HOLTON, J. M..

3 recordsLinked to original sources

WaterFlow: Prediction of Ordered Water Molecule Positions on Protein Structures

Ordered water molecules mediate many protein functions, including stability, ligand binding, and catalysis. Predicting their positions with sub-angstrom accuracy would support protein design, binding affinity prediction, and automated model building in X-ray crystallography and cryo-EM. However, water molecule prediction lags behind protein and other molecule structure predictions. Here, we introduce WaterFlow, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures. WaterFlow outperforms the existing state of the art at every precision level. We demonstrate that WaterFlow can accurately predict ground truth modeled water molecules, including those around protein-ligand interactions and on predicted structures. We also show that WaterFlow predictions fit well directly to experimental data, and therefore propose that it may be used for both prediction and modeling water molecules. This includes novel predictions that are often associated with positive electron difference density, meaning the model places water molecules at sites the original structure depositions omitted. We use this improved model to address the data constraint. By mapping the Pareto front of achievable accuracy of water molecule prediction, alongside analysis of different training data schemas, we quantified the trade-off between data quantity and data quality, demonstrating that the diversity of high-quality structures is limiting the possible results. Overall, WaterFlow predicts ordered water to serve as a solvent module for structure-based drug design and for water molecule placement during crystallographic refinement.

biophysics↗

The Untangle Challenge for accurate ensemble models

We report the discovery of a new class of local minima that has severely limited the accuracy of macromolecular models. Termed density misfit barrier traps, these minima explain much of the poor fit between macromolecular models and experimental data relative to that of smaller molecules: not just high R factors, but distorted chemical geometry. We postulated that proteins exist as an ensemble of conformations that each have good geometry, but refinement algorithms have been unable to converge to them due to a tangling phenomenon arising from these traps. To demonstrate, a synthetic ground truth data set was generated, consisting of a 2-member ensemble with excellent geometry. A series of starting models, each trapped in increasingly difficult local minima, were prepared, a unified validation score defined, and an open Challenge issued. This Challenge inspired algorithms for escaping such traps, and new programs have been released that are expected to substantially improve the accuracy of macromolecular ensemble models. SynopsisA synthetic 2-member conformational ensemble of a small protein and corresponding electron density data was generated to demonstrate how topological local minima hinder simultaneous agreement with density data and chemical geometry restraints in conventional structure refinement.

biophysics↗

Bayesian multi-state multi-condition modeling of a protein structure based on X-ray crystallography data

1An atomic structure model of a protein can be computed from a diffraction pattern of its crystal. While most crystallographic studies produce a single set of atomic coordinates, the billions of protein molecules in a crystal sample many conformational modes during data collection. As a result, a "multi-state" model that depicts these conformations could reproduce the X-ray data better than a single conformation, and thus likely be more accurate. Computing such a multistate model is challenging due to a lower data-to-parameter ratio than that for single-state modeling. To address this challenge, additional information could be considered, such as X-ray datasets collected for the same system under distinct experimental conditions (eg, temperature, ligands, mutations, and pressure). Here, we develop, benchmark, and illustrate MultiXray: Bayesian multi-state multi-condition modeling for X-ray crystallography. The input information is several X-ray datasets collected under distinct conditions and a molecular mechanics force field. The model consists of an independent coordinate set for each of several states and the weight of each state under each condition. A Bayesian posterior model density quantifies the match of the model with all X-ray datasets and the force field. A sample of models is drawn from the posterior model density using biased molecular dynamics (MD) simulations. We benchmark MultiXray on simulated CypA X-ray data. Using a second X-ray dataset improves the Rfree from 0.105 to 0.089. We then demonstrate MultiXray on experimental temperature-dependent data for SARS-CoV-2 Mpro. Using multiple X-ray datasets improves Rfree of the PDB-deposited structure from 0.253 to 0.237. MultiXray is implemented in our open-source Integrative Modeling Platform (IMP) software, relying on integration with Phenix, thus making it easily applicable to many studies.

biophysics↗