bioRxiv Science⌕ Search

Biology subjects

DeBenedictis, E.

Publications and source records attributed to DeBenedictis, E..

4 recordsLinked to original sources

GROQ-seq Datasets Across Transcription Factors (LacI, RamR, VanR), T7 RNA Polymerase and TEV Protease

Predicting any proteins function from its sequence alone would be a significant breakthrough in molecular biology. Although machine learning approaches have sought to tackle this, their limited generalizability reflects the absence of sufficiently large, open, diverse, and unified datasets. To address this data gap, we developed a high-throughput experimental platform called GROQ-seq (Growth-based Quantitative Sequencing). In GROQ-seq, a proteins function can be linked to a sequencing-based readout that enables scalable characterization of large variant libraries in Escherichia coli. Here, we present pilot datasets demonstrating its performance across three distinct protein function classes: transcription factors, polymerases, and proteases. The objective of this report is to present the datasets and to provide users with a clear and transparent characterization of their properties, including both the strengths and limitations.

bioengineering↗

GROQ-seq Enables Cross-site Reproducibility for High-Throughput Measurement of Protein Function

High-throughput functional assays are increasingly used to generate large-scale protein function datasets for protein engineering and machine learning applications. However, the utility of such datasets depends on the reproducibility of the underlying measurements. Here we report reproducible, quantitative measurements of protein sequence-to-function data at scale across two facilities. We analyze GROQ-seq (Growth-based Quantitative Sequencing) measurements of three bacterial transcription factors. Independent barcode measurements of the same sequence produce highly consistent functional estimates, demonstrating strong biological reproducibility (across all transcription factors the mean Root Mean Square Deviation [RMSD] {approx} 0.53 and mean Spearman {approx} 0.63). We also compared experiments performed at two facilities using a shared protocol, but with differing levels of automation and system integration. We observe strong agreement between measurements taken at the two sites (mean RMSD {approx} 0.41 and mean Spearman {approx} 0.730). Orthogonal tests further support this agreement: a classifier trained to distinguish data by site performs near random (AUC = 0.559), and top-ranking variants show strong statistical overlap between experiments. Together, these results demonstrate that GROQ-seq enables reproducible, scalable measurement of protein function suitable for large aggregated datasets.

bioengineering↗

Continuous evolution of a halogenase enzyme with improved solubility and activity for sustainable bioproduction

Halogenation enhances the stability and function of pharmaceuticals, biomaterials, and industrial compounds. However, chemical halogenation lacks stereoselectivity and requires the use of toxic or expensive chemicals. Although enzymatic halogenation can improve selectivity and reduce environmental impact, current halogenases are inefficient and insoluble, leading to low yields that limit their applications. Here, we develop RebHEvo4, a soluble and highly active tryptophan halogenase, containing 12 mutations that confer 37-fold and 44-fold increases in 7-chloro and 7-bromotryptophan production respectively, in vivo. To create RebHEvo4, we devised an aminoacyl tRNA synthetase based halogenase biosensor and conducted over 500 hours of phage-assisted continuous evolution (PACE). Use of RebHEvo4 in a bioreactor resulted in the production of 2.7 g/L of halogenated tryptophan. When coupled with a downstream enzyme, RebHEvo4 allowed 36-fold increased yields of halogenated tryptamines compared to the wild-type enzyme. Additionally, RebHEvo4 enabled efficient production of genetically encoded antimicrobial halogenated peptides. The efficient, site-specific halogenation by our evolved halogenase will accelerate sustainable biomanufacturing of halogenated drugs.

biochemistry↗

Results of the Protein Engineering Tournament: An Open Science Benchmark for Protein Modeling and Design

The grand challenge of protein engineering is the development of computational models to characterize and generate protein sequences for arbitrary functions. Progress is limited by lack of 1) benchmarking opportunities, 2) large protein function datasets, and 3) access to experimental protein characterization. We introduce the Protein Engineering Tournament--a fully-remote competition designed to foster the development and evaluation of computational approaches in protein engineering. The tournament consists of an in silico round, predicting biophysical properties from protein sequences, followed by an in vitro round where novel protein sequences are designed, expressed and characterized using automated methods. Upon completion, all datasets, experimental protocols, and methods are made publicly available. We detail the structure and outcomes of a pilot Tournament involving seven protein design teams, powered by six multi-objective datasets, with experimental characterization by our partner, International Flavors and Fragrances. Forthcoming Protein Engineering Tournaments aim to mobilize the scientific community towards transparent evaluation of progress in the field. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=113 SRC="FIGDIR/small/606135v2_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@8f9769org.highwire.dtl.DTLVardef@11da734org.highwire.dtl.DTLVardef@1cc51c0org.highwire.dtl.DTLVardef@10b3c27_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioengineering↗