bioRxiv Science⌕ Search

Biology subjects

Subramanian, A. M.

Publications and source records attributed to Subramanian, A. M..

4 recordsLinked to original sources

Pretrained protein language models choose between sequence novelty and structural completeness

Protein language models (PLMs) have gained increasing acceptance in tasks ranging from variant effect prediction in disease to optimization and de novo design of proteins with improved stability, target-binding affinity, and catalytic performance. Despite encouraging performance in such applications, little is understood as far as the degree to which PLM-generated sequences - putative novel protein outputs - recapitulate the broad biophysical rules and diversity of sequence, structure, and function that defines natural protein-space, vital knowledge for boosting the design capacity of PLMs in ever-more-complex systems. Towards this end, we computationally profile and characterize the sequence and structure statistics and properties of hundreds of thousands of potential small proteins proposed through free unconstrained generation from architecturally distinct PLMs. We show that although these models exhibit a prodigious latent capacity to access novel amino-acid sequences, they struggle to approach the structural variation that exists on plain display in nature. Moreover, we uncover a stark tradeoff between prioritizing sequence novelty or structural breadth, exemplified by a "helical bundle trap" that dominates model output when aiming outside the comfortable bounds and evolutionary organization of natural sequences. These findings underscore a critical need for strategies that can rapidly guide PLMs into unlocking through generation the full richness of protein sequence, structure, and function that is consistent with governing biophysics but tantalizingly untapped as of yet in design contexts. Author summaryLarge language models (LLMs) like GPT arent just for human text. Spinoff versions that treat protein and DNA sequences as special "languages" of their own, complete with preferred words and grammars, are being used to identify disease-causing genes and mutations and to design new treatments and drugs for clinical testing. But as anyone who has used an LLM chatbot has probably experienced at one time or another, these models can act narrow-minded or nonsensical, reducing their generality and utility. Protein language models are no exception to such flaws of reasoning. We show that protein language models face a fundamental choice between suggesting novel sequences that look nothing like natural ones and capturing the full range of three-dimensional shapes and structures responsible for the diverse functioning of molecular machines. Managing or bypassing this tradeoff is consequently of high import for designing proteins that impart novel functions and activities for therapeutic targeting and beyond.

bioinformatics↗

Rapid discovery of new-to-nature protein domains by novelty-first forcing of language models

Approximations for the existence and extent of physically permissible protein structures beyond those found in nature vary wildly. As predicted structure databases swell thanks to abundant sequence data and generative protein design models concurrently grow in their power to propose new aspects of protein structure, these questions and those of which essential features (e.g. stability, function, robustness) distinguish natural domains from novel ones have been cast in even sharper relief. We demonstrate that protein language models (PLMs) can simultaneously innovate in sequence and structure to suggest new-to-nature protein domains displaying supersecondary and tertiary elements outside of categorized CATH superfamilies. Developing and applying two orthogonal processes for obtaining compact and globular folds from PLMs without bias from other physicochemical or functional constraints, we discover putative novel domains that emerge parallel to known natural ones at rates far exceeding those obtainable by bioinformatic mining of structure databases. Computational characterization of these domain candidates indicates that many exhibit reasonable folding thermodynamics and kinetics, suggesting that natural protein structure-space is far from biophysically complete. These results point away from stability as the definitive selective force behind the observed landscape of real protein folds, and insinuate that many unrealized folds may be equally consistent with the structural rules of protein-based life.

bioengineering↗

Protein CREATE enables closed-loop design of de novo synthetic protein binders

Proteins have proven to be useful agents in a variety of fields, from serving as potent therapeutics to enabling complex catalysis for chemical manufacture. However, they remain difficult to design and are instead typically selected for using extensive screens or directed evolution. Recent developments in protein large language models have enabled fast generation of diverse protein sequences in unexplored regions of protein space predicted to fold into varied structures, bind relevant targets, and catalyze novel reactions. Nevertheless, we lack methods to characterize these proteins experimentally at scale and update generative models based on those results. We describe Protein CREATE (Computational Redesign via an Experiment-Augmented Training Engine), an integrated computational and experimental pipeline that incorporates an experimental workflow leveraging next generation sequencing and phage display with single-molecule readouts to collect vast amounts of quantitative binding data for updating protein large language models. We use Protein CREATE to generate and assay thousands of designed binders to IL-7 receptor and insulin receptor with parallel positive and negative selections to identify on-target binders. We discover not only individual novel binders but also features of ligand-receptor binding, including preservation of the IL7R - ligand hydrophobic interface specifically and existence of multiple approaches to contact the insulin receptor. We also demonstrate the importance of structural features, such as the lack of unpaired cysteine residues, toward design fidelity and find computational pre-screening metrics, such as interchain predicted TM scoring (iPTM), while useful, are imperfect predictors as they neither guarantee experimental binding nor rule it out. We use the data collected from Protein CREATE to score designs from the initial generative models. Globally, Protein CREATE will power future closed-loop design-build-test cycles to enable fine-grained design of protein binders.

bioengineering↗

Unexplored regions of the protein sequence-structure map revealed at scale by a library of foldtuned language models

Amino-acid sequence space is combinatorially vast, with well-folded proteins distributed sparsely and connected by vanishingly few permissible mutational paths. Novel-in-sequence versions of structures observed in nature promise to sample features such as new binding motifs and active site geometries but are rendered inaccessible to evolution or direct search by the extent of sequence perturbations required. Here we introduce a novel algorithm - termed "foldtuning" - that leverages principles of adversarial learning to drive protein language models (PLMs) to erase detectable homology to natural sequences while preserving a target structure, systematically traversing protein-space without being limited by evolutionary barriers. We build foldtuned PLMs for >700 targets including membrane-bound receptors, redox enzymes, and signaling domains. Foldtuned proteins are diverse and far-from-natural in sequence, filling out structurally-equivalent families defined by fundamental biophysical constraints invisible to traditional sequence-based bioinformatics methods. Experimental characterization demonstrates that foldtuned proteins express stably in vitro and function in vivo. By revealing sequence-structure information at scale beyond evolution, foldtuning promises to accelerate the reconstitution and realization of novel-to-nature systems for synthetic biology problems from therapeutics to catalysis.

bioengineering↗