bioRxiv · 10.1101/2022.08.31.505981
From sequence to function through structure: deep learning for protein design
Abstract
The process of designing biomolecules, in particular proteins, is witnessing a rapid change in available tooling and approaches, moving from design through physicochemical force fields, to producing plausible, complex sequences fast via end-to-end differentiable statistical models. To achieve conditional and controllable protein design, researchers at the interface of artificial intelligence and biology leverage advances in natural language processing (NLP) and computer vision techniques, coupled with advances in computing hardware to learn patterns from growing biological databases, curated annotations thereof, or both. Once learned, these patterns can be leveraged to provide novel insights into mechanistic biology and the design of biomolecules. However, navigating and understanding the practical applications for the many recent protein design tools is complex. To facilitate this, we 1) document recent advances in deep learning (DL) assisted protein design from the last three years, 2) present a practical pipeline that allows to go from de novo-generated sequences to their predicted properties and web-powered visualization within minutes, and 3) leverage it to suggest a generated protein sequence which might be used to engineer a biosynthetic gene cluster to produce a molecular glue-like compound. Lastly, we discuss challenges and highlight opportunities for the protein design field. AvailabilitypLM generated and UniRef50 sampled sequence sets and predictions are available at http://data.bioembeddings.com/public/design. Code-base and Notebooks for analysis are available at https://github.com/hefeda/PGP. An online version of Table 1 can be found at https://github.com/hefeda/design_tools. O_TBL View this table: org.highwire.dtl.DTLVardef@1823c4borg.highwire.dtl.DTLVardef@1449504org.highwire.dtl.DTLVardef@193458borg.highwire.dtl.DTLVardef@1bab4c4org.highwire.dtl.DTLVardef@b1d39f_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 1:C_FLOATNO O_TABLECAPTIONmachine-learning-based protein design methods. Methods ordered by their release date, accounting for the date of pre-prints when available. Class (1-4) captured in Fig. 1. Expanded method detail in main text. Unnamed methods are referenced by the first name of the first author. Online version at https://github.com/hefeda/design_tools. Legend: Architecture: The architecture of the deep learning model; E2E: an end-to-end differentiable solution; Input: the input to run (infer from) the model, e.g. a contact map; Output: the output, e.g. a protein sequence; Dataset: the number and type of samples used to train the method; Params: the exact or estimated number of parameters of the model; 1: the 3D recovery was performed in an external, second step; 2: conditioned generation; -: no input required for generation. C_TABLECAPTION C_TBL
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ferruz, N., Heinzinger, M., Akdel, M., Goncearenco, A., Naef, L., Dallago, C.. 2022-09-03. From sequence to function through structure: deep learning for protein design. https://doi.org/10.1101/2022.08.31.505981
Cite the original work for its findings. Save a collection to share your selection of sources.