bioRxiv Science⌕ Search

Biology subjects

Belay, F.

Publications and source records attributed to Belay, F..

3 recordsLinked to original sources

Evaluating codon optimization strategies for mammalian glycoprotein production with an open-source expression vector

Efficient production of human proteins for the development of tool compounds and biologics depends on a detailed understanding of the protein expression machinery in mammalian cells. Codon optimization is widely believed to enhance protein yield, yet its impact in homologous mammalian systems remains poorly defined. Here, we systematically compare five codon usage strategies reflecting common assumptions about rare codons, RNA stability, and synthesis efficiency. We developed pTipi, an efficient open-source mammalian expression vector, and evaluated its performance in antibody production. We generated plasmids for common epitope tag antibodies such as V5, anti-biotin and anti-His for distribution by Addgene. To compare codon usage schemes, we performed a bake-off of 18 human and murine Wnt pathway glycoproteins in mammalian cells. Small-scale expression screens revealed that codon optimization did not provide a general advantage over native coding sequences, while strategies prioritizing RNA stability consistently reduced expression. Interestingly, a skewed codon scheme using the most abundant codons produced yields comparable to native sequences and occasionally enhanced protein output. To enable flexible evaluation of codon strategies, we implemented a Golden Gate-compatible pTipi platform for efficient synthetic gene incorporation. We conclude that native codons are sufficient for robust homologous mammalian expression of glycoproteins, while selective codon skewing can be beneficial for some targets.

molecular biology↗

Machine Learning enables efficient and effective affinity maturation of nanobodies

Antibodies can bind their targets with exquisite potency and selectivity due in part to large antibody-target protein-protein interaction surface areas. Despite the very large size and diversity of synthetic libraries, in vitro sorting alone tends to yield binders with modest affinities. By analogy to the in vivo affinity maturation in the natural immune system, these initial hits are typically affinity matured in vitro to achieve high affinity binding. However, affinity maturation campaigns can be laborious, often requiring multiple selection rounds and strategies for each clone to be optimized. Here, we investigated whether one could accelerate the discovery of optimized binders using machine learning on sequencing data from single selection sorts of affinity maturation yeast-display campaigns. Our results show that sparse sequencing data from a single sorting round can predict sequences that are enriched after multiple rounds. We also find that linear models outperform deep neural networks and semi-supervised approaches in ranking validated affinity-enhancing substitutions. Linear models are also more interpretable, offering insights into residue preferences that can be leveraged for further engineering. We use our models to design and select optimized nanobody binders to relaxin family peptide receptor 1 (RXFP1), yielding multiple improved binders including 3 sub nanomolar binders with the best exhibiting a [~]2500-fold improvement over WT.

bioinformatics↗

High-Throughput Machine Learning-Aided Antibody Discovery for Cell Surface Antigens

Machine learning (ML) has the potential to revolutionize antibody design and selection, but its success depends on access to extensive, well-curated datasets of antibody-antigen interactions. To address this need, we developed a synthetic Fab yeast display library optimized for seamless ML integration, focusing on sequence diversity within the CDRH3 loop. The library incorporates key sequence features derived from human B cell repertoires essential for efficient antibody generation captured in a compact antigen recognition module (ARM) format. Built using the VH1-69 heavy chain and four light chains, the library was evaluated against ten human and murine cell surface antigens, including PD-L1, TIGIT, and ROBO1. This approach yielded hundreds of antibodies with robust biophysical properties, validated for functional performance in flow cytometry and immunohistochemistry. Furthermore, ML analysis identified additional antibodies for ROBO2 and PD-L2 from the aggregate sequencing data, demonstrating utility for hybrid in silico and experimental workflows. We provide a publicly accessible dataset comprising more than 68,000 Fab sequences and 486 characterized antibodies. This study establishes an ML-compatible framework designed to accelerate and streamline antibody discovery and development.

biophysics↗