bioRxiv Science⌕ Search

Biology subjects

Freschlin, C.

Publications and source records attributed to Freschlin, C..

2 recordsLinked to original sources

Scalable and cost-efficient custom gene library assembly from oligopools

Advances in metagenomics, deep learning, and generative protein design have enabled broad in silico exploration of sequence space, but experimental characterization is still constrained by the cost and scalability of DNA synthesis. Here, we present OMEGA (Oligo-based Multiplexed Efficient Gene Assembly), a low-cost, accessible method for assembling hundreds to thousands of full-length genes in parallel using standard laboratory techniques. OMEGA computationally fragments target genes into short, high-fidelity Golden Gate-compatible oligonucleotides that can be ordered as a pooled library and assembled across multiplexed subpools. We systematically optimized the number of fragments per gene and orthogonal ligation sites per reaction and determine that OMEGA can assemble up to 2.6 kb constructs using as many as 70 Golden Gate sites. To validate the approach, we assembled and functionally screened a library of 810 natural and synthetic GFP variants, recovering 94-97% of target sequences with high uniformity. OMEGA enables precision library construction at scale, with per-gene costs as low as $1.50, and offers a broadly applicable solution for bridging computational protein design with high-throughput experimental validation. We have developed OMEGA as an open-source software package and an easy-to-use Colab notebook available at https://github.com/RomeroLab/omega.

synthetic biology↗

Biophysics-based protein language models for protein engineering

Protein language models trained on evolutionary data have emerged as powerful tools for predictive problems involving protein sequence, structure, and function. However, these models overlook decades of research into biophysical factors governing protein function. We propose Mutational Effect Transfer Learning (METL), a protein language model framework that unites advanced machine learning and biophysical modeling. Using the METL framework, we pretrain transformer-based neural networks on biophysical simulation data to capture fundamental relationships between protein sequence, structure, and energetics. We finetune METL on experimental sequence-function data to harness these biophysical signals and apply them when predicting protein properties like thermostability, catalytic activity, and fluorescence. METL excels in challenging protein engineering tasks like generalizing from small training sets and position extrapolation, although existing methods that train on evolutionary signals remain powerful for many types of experimental assays. We demonstrate METLs ability to design functional green fluorescent protein variants when trained on only 64 examples, showcasing the potential of biophysics-based protein language models for protein engineering.

bioinformatics↗