bioRxiv Science⌕ Search

Biology subjects

Penner, M.

Publications and source records attributed to Penner, M..

7 recordsLinked to original sources

Latent generative search unlocks de novo design of untapped biomolecular interactions at scale

De novo protein design has advanced rapidly, yet designing binders to polar, solvent-exposed epitopes and small, flexible ligands remains challenging. Such hydrated surfaces and flexible molecules, including carbohydrates, provide few of the hydrophobic contacts favoured by current methods and have largely resisted de novo binders. To address this challenge, here we introduce latent generative search for binder design, a novel framework that uses reward-guided search at inference time to steer the Proteina-Complexa generative model. The model codesigns sequence and structure - generating them together in a continuous latent space - and thereby removes the inverse-folding step on which current methods rely. In a screen of more than one million designs by multiplexed phage display, latent generative search produced more validated binders than every other method tested, its codesigned sequences surpassing post hoc redesign. It delivered high-affinity binders across therapeutic receptors, a viral attachment protein and intracellular signalling targets. Our approach also accessed previously untapped biology, generating the first de novo proteins that bind a free carbohydrate, including one that discriminates between blood-group antigens - a polar, flexible target class beyond the reach of current design methods.

bioengineering↗

Unsupervised protein language models learn patterns of enzyme function

While enormous amounts of sequence information have become available, assignment of sequence to a particular enzymatic function has remained elusive. Here we describe a framework that drives a general protein language model to find a target reaction without specific training, using an initial bridgehead protein. At the heart of this framework is PLM-clust, an algorithm that employs k-means on top of protein language model embeddings to convert sequence space into functional reservoirs of latent space, and samples from these clusters based on accelerated zero-shot scoring. We demonstrate PLM-clust in a recursive discovery process (with enzyme hit rates quickly rising to >90%), segmenting isofunctional reservoirs and exploring them in greater detail. This approach - exemplified for glycosyl hydrolases (a xylanase, >100-fold activity increase) and for imine reductases (IREDs, >100-fold increase in catalytic promiscuity profiles) - reliably brings about novel enzymes that are proficient at the catalytic task at hand, reaching deeply into sequence space with a majority of residues exchanged.

synthetic biology↗

A metagenomic thermostable monomeric meganuclease with novel specificity and unique palindromic 3-prime overhangs

As an alternative to historical enzyme isolation, metagenomic databases (e.g. MGnify) provide information on vast unculturable microbial diversity, especially from extreme environments, and constitute an enormous source of functional proteins. Conservative mining of these data by close sequence homology alone tends to identify merely different versions of known enzymes. Here we present a discovery strategy of meganucleases based on wider capture of less homologous enzymes with new function in metagenomic databases, incorporating metadata with homology, relying on cell-free expression to bypass host incompatibility and the need for purification, along with using deep sequencing for experimental assessment of substrate specificity and cleavage pattern, circumventing classical gel-based profiling. Specifically, we discovered the temperature-stable (>55{degrees}C), intron-encoded LAGLIDADG meganuclease I-MG11 that recognizes a 17 base pair sequence to generate unique 4 base pair palindromic 3'-overhangs -- the first monomeric meganuclease to produce such overhangs. Co-folding models of I-MG11 bound to DNA provide a structural context for enzyme-DNA interactions, highlighting differences from other monomeric LAGLIDADG meganucleases (e.g. I-SceI) shaped by InDels (insertion-deletions) in the DNA binding region that may cause specificity changes. Our strategy streamlines bona fide identification and annotation of meganucleases, while the unique properties of I-MG11 expand the molecular biology toolbox. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=31 SRC="FIGDIR/small/712669v1_ufig1.gif" ALT="Figure 1"> View larger version (11K): org.highwire.dtl.DTLVardef@199394aorg.highwire.dtl.DTLVardef@806490org.highwire.dtl.DTLVardef@14a533aorg.highwire.dtl.DTLVardef@9e13db_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry↗

Convergent Acquisition of Glucomannan β-galactosyltransferases in Asterids and Rosids

{beta}-Galactoglucomannan ({beta}-GGM) is a primary cell wall polysaccharide in rosids and asterids. The {beta}-GGM polymer has a backbone of repeating glucose and mannose, usually with mono- or di-galactosyl sidechains on the mannosyl residues. CELLULOSE SYNTHASE-LIKE 2 (CSLA2), MANNAN -GALACTOSYLTRANSFERASE (MAGT), and MANNAN {beta}-GALACTOSYLTRANSFERASE (MBGT) are required for {beta}-GGM synthesis in Arabidopsis thaliana. The single MBGT identified so far, AtMBGT1, lies in glycosyltransferase family 47A subclade VII, and was identified in Arabidopsis. However, despite the presence of {beta}-GGM, an orthologous gene is absent in tomato (Solanum lycopersicum), a model asterid. In this study, we screened candidate MBGT genes from the tomato genome, functionally tested the activities of encoded proteins, and identified the tomato MBGT (SlMBGT1) in GT47A-III. Interestingly therefore, AtMBGT1 and SlMBGT1 are located in different GT47A subclades. Further, phylogenetic and glucomannan structural analysis from different species raised the possibility that various asterids possess conserved MBGTs in GT47A-III, indicating that MBGT activity has been acquired convergently among asterids and rosids. Although functional convergence was observed, the acquired amino acid substitutions among the two MBGT groups were not shared, suggesting different evolutionary pathways to achieve the same biochemical outcome. The present study highlights the promiscuous emergence of donor and acceptor preference in GT47A enzymes, and suggests an adaptive advantage for eudicots to acquire {beta}-GGM {beta}-galactosylation.

plant biology↗

Sub-single-turnover quantification of enzyme catalysis at ultrahigh throughput via a versatile NAD(P)H coupled assay in microdroplets

Enzyme engineering and discovery are crucial for a future sustainable bioeconomy. Harvesting new biocatalysts from large libraries through directed evolution or functional metagenomics requires accessible, rapid assays. Ultra-high throughput screening formats often require optical readouts, leading to the use of model substrates that may misreport target activity and necessitate bespoke synthesis. This is a particular challenge when screening glycosyl hydrolases, which leverage molecular recognition beyond the target glycosidic bond, so that complex chemical synthesis would have to be deployed to build a fluoro- or chromogenic substrate. In contrast, coupled assays represent a modular plug-and-play system: any enzyme- substrate pairing can be investigated, provided the reaction can produce a common intermediate which links the catalytic reaction to a detection cascade readout. Here, we establish a detection cascade producing a fluorescent readout in response to NAD(P)H via glutathione reductase and a subsequent thiol-mediated uncaging reaction, with a low nanomolar detection limit in plates. Further scaling down to microfluidic droplet screening is possible: the fluorophore is leakage- free and we report a three orders of magnitude improved sensitivity compared to absorbance- based systems, so that less than one turnover per enzyme molecule expressed from a single cell is detectable. Our approach enables the use of non-fluorogenic substrates in droplet-based enrichments, with applicability in screening for glycosyl hydrolases and imine reductases (IREDs). To demonstrate the assays readiness for combinatorial experiments, one round of directed evolution was performed to select a glycosidase processing a natural substrate, beechwood xylan, with improved kinetic parameters from a pool of >106 mutagenized sequences.

biochemistry↗

YeastIT: Reducing mutational bias for in vivo directed evolution using a novel yeast mutator strain based on dual adenine-/cytosine-targeting and error-prone DNA repair

Engineering proteins with new functions and properties often requires navigating large sequence spaces through rounds of iterative improvement. However, a disparity exists between the gradual pace of natural long-term evolution and a typical laboratory evolution workflow that relies on enriching functional variants from highly diverse in vitro generated libraries through very few screening rounds. Laboratory experiments often eschew presumed natural strategies such as neutral/non-adaptive and multi-phase evolution trajectories, and therefore mutagenesis technologies suitable for long nature-like timescales are needed. Here, we introduce YeastIT, a novel in vivo mutagenesis tool for protein engineering that leverages an S. cerevisiae strain engineered to exhibit mutagenic activity directed to the gene of interest, allowing its continuous diversification. Mutagenesis is achieved by generating DNA damage through nucleoside deamination, followed by introduction of mutations by harnessing the process of error-prone DNA translesion synthesis. By eliminating the transformation step, YeastIT allows multiple rounds of screening or selection without interruptions for library diversification, thereby enabling long-term and continuous evolution campaigns. Our characterization of the mutational spectrum and frequency of the YeastIT-generated libraries, and its comparison to other methods (error-prone PCR, PACE, MutaT7, eMutaT7, OrthoRep, TRIDENT, EvolVR) demonstrates comparable mutation rates combined with a significant reduction in mutagenic bias relative to most of the alternatives. To validate YeastIT, we carried out directed evolution of a DARPin binding protein to achieve a 15-fold improved affinity. YeastIT thus provides a tool for exploring different evolutionary trajectories which overcomes previous limitations of variant availability (due to bias and low mutation rates) and emulates the way proteins emerge in Nature.

synthetic biology↗

Altered sleep intensity upon DBS to hypothalamic sleep-wake centers in rats

Deep brain stimulation (DBS) has been scarcely investigated in the field of sleep research. We hypothesize that DBS onto hypothalamic sleep- and wake-promoting centers will produce significant neuromodulatory effects, and potentially become a therapeutic strategy for patients suffering severe, drug-refractory sleep-wake disturbances. We aimed to investigate whether continuous electrical high-frequency DBS, such as that often implemented in clinical practice, in the ventrolateral preoptic nucleus (VLPO) or the perifornical area of the posterior lateral hypothalamus (PeFLH), significantly modulates sleep-wake characteristics and behavior. We implanted healthy rats with electroencephalographic/electromyographic electrodes and recorded vigilance states in parallel to bilateral bipolar stimulation of VLPO and PeFLH at 125 Hz at 90 A over 24 h to test the modulating effects of DBS on sleep-wake proportions, stability and spectral power in relation to baseline. We unexpectedly found that VLPO DBS at 125 Hz deepens slow-wave sleep as measured by increased delta power, while sleep proportions and fragmentation remain unaffected. Thus, the intensity, but not the amount of sleep or its stability, is modulated. Similarly, the proportion and stability of vigilance states remained altogether unaltered upon PeFLH DBS but, in contrast to VLPO, 125 Hz stimulation unexpectedly weakened SWS, evidenced by reduced delta power. This study provides novel insights into non-acute functional outputs of major sleep-wake centers in the rat brain in response to electrical high-frequency stimulation, a paradigm frequently used in human DBS. In the conditions assayed, while exerting no major effects on sleep-wake architecture, hypothalamic high-frequency stimulation arises as a provocative sleep intensity-modulating approach.

neuroscience↗