bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.02.14.705946

Universal Baseline for in vitro Selection of Genetically Encoded Libraries

Abstract

Genetically encoded (GE) libraries enable identification of high-affinity ligands for diverse molecular targets through iterative in vitro selection and DNA sequencing or next-generation sequencing (NGS). Despite their impact in therapeutic development, a systematic framework for evaluating reproducibility in GE-molecular discoveries remains limited. To aid such analysis, we introduce the concept of baseline response, which reproducibly partitions active and inactive members of in vitro selection. The baseline response is provided by spiking a random DNA-barcoded population. We calibrated the baseline concept using Bioconductor EdgeR differential enrichment (DE) analysis of NGS of phage-displayed selection on oligosaccharide chitin and hepatitis virus NS3a* protease as model targets. We further show that mixing discovery campaigns also offers an effective baseline: chitin-enriched peptides serve as a baseline for DE-analysis of NS3a* selection and NS3a*-enriched peptides serve as a baseline for chitin binders. We applied baseline-stratified DE-analysis to 66 parallel selections performed in 3-5 replicates across 22 extracellular targets, including HER1-3, EpCAM, CAIX, PD-L1, and eight integrin receptors. Automated DE-analysis across hundreds of NGS files produced hits validated in a secondary screen and yielded synthetic macrocyclic ligands with mid-nanomolar affinity confirmed in 2-3 biophysical assays. For PD-L1, we further demonstrated how baseline-calibrated NGS data provide decision-enabling information for optimization of peptide macrocycles to yield potent single-digit nanomolar ligands for the cell-surface receptor. We anticipate that baseline-based analyses of NGS data from in vitro selection procedures will offer a scalable framework for reproducible hit discovery and standardized analysis across diverse in vitro selection campaigns. Significance StatementGenetically encoded selection technologies such as phage, mRNA and ribosome display, have produced FDA-approved therapeutics and numerous clinical candidates. Yet reproducibility in such in vitro discovery systems is rarely evaluated against a defined experimental baseline. Here, we establish a universal baseline by spiking unrelated, DNA-barcoded peptide sequences into selection libraries and quantifying their binding alongside target-enriched populations. This composition-agnostic strategy enables rigorous normalization, confidence assessment, and cross-target comparison of molecular discovery outcomes. Our framework introduces practical standards for reproducibility and statistical benchmarking across genetically encoded display platforms.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yan, K., Lima, G. M., Bahadur, T., Albert, V., O'Gara, Z., Bao, G., Kossmann, C., Kirby, W., Mejia, F. B., Michnik, M. L., Maiorana, K., Derda, R.. 2026-02-15. Universal Baseline for in vitro Selection of Genetically Encoded Libraries. https://doi.org/10.64898/2026.02.14.705946

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

aaRSID, an engineered pyrrolysyl-tRNA synthetase platform for multi-probe proximity proteomics

Proximity labeling (PL) methods utilize spatially targeted chemical or enzymatic generation of a diffusible, reactive intermediate to covalently tag neighboring proteins in living systems. Unlike other tools for studying molecular interactions, PL can detect transient protein relationships with high spatial and temporal sensitivity, allowing for insight into their roles in biological processes. However, current enzymatic PL tools, such as TurboID and APEX2, are limited by their substrate structure and chemistry, which can generate significant background and/or perturb cellular physiology. To address these limitations, we have developed aminoacyl-tRNA synthetase ID (aaRSID), a PL tool that leverages an engineered pyrrolysyl tRNA synthetase (PylRS) for proximity labeling of proteins. We chose PylRS because it can catalyze promiscuous lysine labeling in the absence of its cognate tRNA and utilize a variety of non-canonical amino acids (ncAAs) as substrates. Here, we demonstrate aaRSID's intrinsic proximity labeling activity, use directed evolution to improve this activity, and apply the improved mutant (aaRSID-Ma1.3) for subcellular proteomics and multiplexed imaging. Our work establishes aminoacyl-tRNA synthetases as a new PL enzyme class and introduces a versatile chemical platform for developing ncAA-derived probes to map cellular microenvironments, greatly expanding the applications possible of PL technology.

biochemistry↗

Cellular uptake of folate-olaparib conjugates via folate receptor-mediated endocytosis: Potential for selective delivery of DNA damage response inhibitors into tumour cells

The folate receptor (FR) is overexpressed in a range of human tumours including ovarian cancer cells. We propose that the overexpression of the FR on the surface of ovarian tumour cells could be exploited for the selective delivery of a DNA damage response inhibitor (DDRi) in the form of an intact folate drug conjugate (FDC). This approach would improve the therapeutic index of the parent DDRi facilitating combination studies of the DDRi-based FDC with DNA damaging chemotherapy. FR-mediated cellular uptake of the proposed folate drug conjugates is requisite for FDC selective delivery into tumours. In this study, we synthesised a series of olaparib-based folate conjugates that maintained the biochemical PARP1 inhibition associated with olaparib and showed binding affinity for the folate receptor. Significantly, we identified compounds 10b and 11 that selectively enter FR overexpressing tumour cells via folate receptor-mediated endocytosis in their intact form and engage with their target as demonstrated by the potent inhibition of PARylation (KB cells, PARylation IC50 = 5.7 and 3.9 nM; respectively).

biochemistry↗

Architecture and Energy Transfer of the Bacterial Photosynthetic Unit

In phototrophic organisms, pigment-protein membrane complexes are densely packed to form photosynthetic units (PSUs) that capture solar energy and convert it into chemical energy. Although the structures of many individual photosynthetic complexes have been resolved, how they are arranged and interact with others within photosynthetic membranes to enable efficient excitation energy transfer (EET) remains poorly understood. Here, we report cryo-electron microscopy structures of PSU supercomplex assemblies from the phototrophic a-proteobacterium Rhodovulum viride, including an RC-LH1 core associated with one or two peripheral LH2 complexes and a curved LH2 tetramer. These membrane-derived assemblies define the relative positions and orientations of neighboring photosynthetic complexes and place their pigment arrays in proximity across antenna-antenna and antenna-core interfaces. Structure-based simulations identify potential EET pathways within the PSU assemblies and reveal rapid energy transfer across both LH2-LH2 and LH2-LH1 interfaces. Collectively, these findings provide insights into the assembly and structural modularity of bacterial PSUs and elucidate how the lateral organization of membrane protein complexes facilitates efficient energy transfer. This work extends structural studies of bacterial photosynthesis from individual complexes to their native higher-order assembly, providing a framework for understanding how photosynthetic supercomplex organization shapes energy migration and for guiding the design of artificial photosynthesis.

biochemistry↗