bioRxiv Science⌕ Search

Biology subjects

Ugarte La Torre, D.

Publications and source records attributed to Ugarte La Torre, D..

2 recordsLinked to original sources

CGBack: Diffusion Model for Backmapping Large-Scale and Complex Coarse-Grained Molecular Systems

Molecular dynamics simulations based on coarse-grained (CG) models are used to accelerate conformational dynamics of biomolecules and other chemical systems with reduced computational costs. CG models achieve it by discarding atomic information necessary for downstream structural analysis. Recovering the atomic detail from CG structures, i.e. backmapping, remains a fundamental challenge in multiscale modeling, especially for proteins and complex biomolecular assemblies. Recent machine learning methods have shown promise in reconstructing atomistic details, but most approaches only target simple systems at a small scale. In particular, conventional backmapping pipelines often fail to preserve stereochemistry, induce high-energy configurations, or require extensive minimization. Here we present CGBack, a backmapping framework that employs a denoising diffusion probabilistic model to reconstruct all-atom molecular structures from CG representations. CGBack consists of backmapping and refinement procedures and accurately recovers atomic coordinates across diverse protein systems from small to large scales. We show that CGBack is capable of accurately backmapping both single-chain and multi-chain molecular systems, including densely packed intrinsically disordered proteins in condensates. These results suggest that CGBack can be a powerful tool for multiscale molecular simulation pipelines. We anticipate that CGBack will enable more efficient workflows for protein modelling, as well as for other biomolecules across different CG models.

biophysics↗

Implementation of residue-level coarse-grained models in GENESIS for large-scale molecular dynamics simulations

Residue-level coarse-grained (CG) models have become one of the most popular tools in biomolecular simulations in the trade-off between modeling accuracy and computational efficiency. To investigate large-scale biological phenomena in molecular dynamics (MD) simulations with CG models, unified treatments of proteins and nucleic acids, as well as efficient parallel computations, are indispensable. In the GENESIS MD software, we implement several residue-level CG models, covering structure-based and context-based potentials for both well-folded biomolecules and intrinsically disordered regions. An amino acid residue in protein is represented as a single CG particle centered at the C atom position, while a nucleotide in RNA or DNA is modeled with three beads. Then, a single CG particle represents around ten heavy atoms in both proteins and nucleic acids. The input data in CG MD simulations are treated as GROMACS-style input files generated from a newly developed toolbox, GENESIS-CG-tool. To optimize the performance in CG MD simulations, we utilize multiple neighbor lists, each of which is attached to a different nonbonded interaction potential in the cell-linked list method. We found that random number generations for Gaussian distributions in the Langevin thermostat are one of the bottlenecks in CG MD simulations. Therefore, we parallelize the computations with message-passing-interface (MPI) to improve the performance on PC clusters or supercomputers. We simulate Herpes simplex virus (HSV) type 2 B-capsid and chromatin models containing more than 1,000 nucleosomes in GENESIS as examples of large-scale biomolecular simulations with residue-level CG models. This framework extends accessible spatial and temporal scales by multi-scale simulations to study biologically relevant phenomena, such as genome-scale chromatin folding or phase-separated membrane-less condensations. Author summaryMolecular dynamics (MD) simulations have been widely used to investigate biological phenomena that are difficult to study only with experiments. Since all-atom MD simulations of large biomolecular complexes are computationally expensive, coarse-grained (CG) models based on different approximations and interaction potentials have been developed so far. There are two practical issues in biological MD simulations with CG models. The first issue is the input file generations of highly heterogeneous systems. In contrast to well-established all-atom models, specific features are introduced in each CG model, making it difficult to generate input data for the systems containing different types of biomolecules. The second issue is how to improve the computational performance in CG MD simulations of heterogeneous biological systems. Here, we introduce a user-friendly toolbox to generate input files of residue-level CG models containing folded and disordered proteins, RNAs, and DNAs using a unified format and optimize the performance of CG MD simulations via efficient parallelization in GENESIS software. Our implementation will serve as a framework to develop novel CG models and investigate various biological phenomena in the cell.

biophysics↗