bioRxiv Science⌕ Search

Biology subjects

McHugh, L.

Publications and source records attributed to McHugh, L..

4 recordsLinked to original sources

RNA sequence design and protein-DNA specificity prediction with NA-MPNN

RNA sequence design and protein-DNA binding specificity prediction can both be framed as nucleic acid inverse-folding problems: finding the most likely nucleic acid sequences given a fixed three-dimensional structure of a nucleic acid or nucleic acid-protein complex. While task-specific tools have been developed, no unified deep learning model for nucleic acid inverse folding has been described; a single model would have larger and more diverse datasets available for training and a considerably greater range of applicability. Here we introduce Nucleic Acid MPNN (NA-MPNN), a message-passing neural network that treats proteins, DNA, and RNA within a unified biopolymer graph representation. NA-MPNN outperforms previous methods on RNA sequence design and fixed-dock protein-DNA specificity prediction, and should be broadly useful for de novo RNA structure design and prediction of DNA-binding specificity.

biochemistry↗

De novo design of RNA and nucleoprotein complexes

Nucleic acids fold into sequence-dependent tertiary structures and carry out diverse biological functions, much like proteins. However, while considerable advances have been made in the de novo design of protein structure and function, the same has not yet been achieved for RNA tertiary structures of similar intricacy. Here, we describe a generative diffusion framework, RFDpoly, for generalized de novo biopolymer (RNA, DNA and protein) design, and use it to create diverse and designable RNA structures. We design RNA structures with novel folds and experimentally validate them using a combination of chemical footprinting (SHAPE-seq) and electron microscopy. We further use this approach to design protein-nucleic acid assemblies; the crystal structure of one such design is nearly identical to the design model. This work demonstrates that the principles of structure-based de novo protein design can be extended to nucleic acids, opening the door to creating a wide range of new RNA structures and protein-nucleic acid complexes.

biochemistry↗

Accelerating Biomolecular Modeling with AtomWorks andRF3

Deep learning methods trained on protein structure databases have revolutionized biomolecular structure prediction, but developing and training new models remains a considerable challenge. To facilitate the development of new models, we present AtomWorks: a broadly applicable data framework for developing state-of-the-art biomolecular foundation models spanning diverse tasks, including structure prediction, generative protein design, and fixed backbone sequence design. We use AtomWorks to train RosettaFold-3 (RF3), a structure prediction network capable of predicting arbitrary biomolecular complexes with an improved treatment of chirality that narrows the performance gap between closed-source AlphaFold3 (AF3) and existing open-source implementations. We expect that AtomWorks will accelerate the next generation of open-source biomolecular machine learning models and that RF3 will be broadly useful as a structure prediction tool. To this end, we release the AtomWorks framework (https://github.com/RosettaCommons/atomworks), together with curated training data, code and model weights for RF3 (https://github.com/RosettaCommons/modelforge) under a permissive BSD license.

biochemistry↗

The MicroMap is a network visualisation resource for microbiome metabolism

The human microbiome plays a crucial role in metabolism and thereby influences health and disease. Constraint-based reconstruction and analysis (COBRA) has proven an attractive framework to generate mechanism-derived hypotheses along the nutrition-host-microbiome-disease axis within the computational systems biology community. Unlike for human, no large-scale visualisation resource for microbiome metabolism has been available to date. To address this gap, we created the MicroMap, a manually curated microbiome metabolic network visualisation, which captures the metabolic content of over a quarter million microbial genome-scale metabolic reconstructions. The MicroMap contains 5,064 unique reactions and 3,499 unique metabolites, including for 98 drugs. The MicroMap allows users to intuitively explore microbiome metabolism, inspect microbial metabolic capabilities, and visualise computational modelling results. Further, the MicroMap shall serve as an educational tool to make microbiome metabolism accessible to broader audiences beyond computational modellers. For example, we utilised the MicroMap to generate a comprehensive collection of 257,429 visualisations, corresponding to the entire scope of our current microbiome reconstruction resources, to enable users to visually compare and contrast the metabolic capabilities for diaerent microbes. The MicroMap seamlessly integrates with the Virtual Metabolic Human (VMH, www.vmh.life) and the COBRA Toolbox (opencobra.github.io), and is freely accessible at the MicroMap dataverse (https://dataverse.harvard.edu/dataverse/micromap), in addition to all the generated reconstruction visualisations.

systems biology↗