bioRxiv · 10.1101/2023.12.22.573145
Unexplored regions of the protein sequence-structure map revealed at scale by a library of foldtuned language models
Abstract
Amino-acid sequence space is combinatorially vast, with well-folded proteins distributed sparsely and connected by vanishingly few permissible mutational paths. Novel-in-sequence versions of structures observed in nature promise to sample features such as new binding motifs and active site geometries but are rendered inaccessible to evolution or direct search by the extent of sequence perturbations required. Here we introduce a novel algorithm - termed "foldtuning" - that leverages principles of adversarial learning to drive protein language models (PLMs) to erase detectable homology to natural sequences while preserving a target structure, systematically traversing protein-space without being limited by evolutionary barriers. We build foldtuned PLMs for >700 targets including membrane-bound receptors, redox enzymes, and signaling domains. Foldtuned proteins are diverse and far-from-natural in sequence, filling out structurally-equivalent families defined by fundamental biophysical constraints invisible to traditional sequence-based bioinformatics methods. Experimental characterization demonstrates that foldtuned proteins express stably in vitro and function in vivo. By revealing sequence-structure information at scale beyond evolution, foldtuning promises to accelerate the reconstitution and realization of novel-to-nature systems for synthetic biology problems from therapeutics to catalysis.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Subramanian, A. M., Thomson, M.. 2023-12-23. Unexplored regions of the protein sequence-structure map revealed at scale by a library of foldtuned language models. https://doi.org/10.1101/2023.12.22.573145
Cite the original work for its findings. Save a collection to share your selection of sources.