bioRxiv Science⌕ Search

Biology subjects

McPartlon, m.

Publications and source records attributed to McPartlon, m..

2 recordsLinked to original sources

End-to-End deep structure generative model for protein design

AO_SCPLOWBSTRACTC_SCPLOWDesigning protein with desirable structure and functional properties is the pinnacle of computational protein design with unlimited potentials in the scientific community from therapeutic development to combating the global climate crisis. However, designing protein macromolecules at scale remains challenging due to hard-to-realize structures and low sequence design success rate. Recently, many generative models are proposed for protein design but they come with many limitations. Here, we present a VAE-based universal protein structure generative model that can model proteins in a large fold space and generate high-quality realistic 3-dimensional protein structures. We illustrate how our model can enable robust and efficient protein design pipelines with generated conformational decoys that bridge the gap in designing structure conforming sequences. Specifically, sequences generated from our design pipeline outperform native fixed backbone design in 856 out of the 1,016 tested targets(84.3%) through AF2 validation. We also demonstrate our models design capability and structural pre-training potential by structurally inpainting the complementarity-determining regions(CDRs) in a set of monoclonal antibodies and achieving superior performance compared to existing methods.

bioinformatics↗

A Deep SE(3)-Equivariant Model for Learning Inverse Protein Folding

In this work, we establish a framework to tackle the inverse protein design problem; the task of predicting a proteins primary sequence given its backbone conformation. To this end, we develop a generative SE(3)-equivariant model which significantly improves upon existing autoregressive methods. Conditioned on backbone structure, and trained with our novel partial masking scheme and side-chain conformation loss, we achieve state-of-the-art native sequence recovery on structurally independent CASP13, CASP14, CATH4.2, and TS50 test sets. On top of accurately recovering native sequences, we demonstrate that our model captures functional aspects of the underlying protein by accurately predicting the effects of point mutations through testing on Deep Mutational Scanning datasets. We further verify the efficacy of our approach by comparing with recently proposed inverse protein folding methods and by rigorous ablation studies.

bioinformatics↗