bioRxiv Science⌕ Search

Biology subjects

Cross, R. S.

Publications and source records attributed to Cross, R. S..

2 recordsLinked to original sources

Active Learning Enables Efficient Directed Evolution of a Far-Red Fluorescent Protein with Minimal Experimental Data

Fluorescent proteins are fundamental tools for cellular imaging. Most fluorescent proteins in routine use, including GFP, are derived from the jellyfish Aequorea victoria and emit blue-green light, which is strongly absorbed and scattered by tissue, limiting imaging depth. Far-red and near-infrared fluorescent proteins, engineered from bacteriophytochromes, address this limitation because far-red light penetrates tissue considerably further. However, these proteins are typically much dimmer than their A. victoria -derived counterparts. Improving brightness by conventional directed evolution requires screening large random mutant libraries, a process that is slow, labor-intensive, and often impractical outside specialized laboratories. We utilized an active-learning-guided directed evolution workflow that identified improved variants from substantially less data than conventional screening. Each round coupled automated, miniaturized cell-free protein expression directly from a DNA template without cloning or cell culture, with a machine-learning model retrained on cumulative sequence-function data to nominate the most informative variants for the next round. Applied to miRFP670nano3, this workflow screened 120 variants across successive rounds and identified twelve with improved brightness, the best four-fold brighter in bacterial systems. However, these gains did not translate when the variants were evaluated in mammalian cells, indicating that performance can be strongly dependent on cellular context. Retrospective simulation across benchmark datasets from ProteinGym showed that performing more experimental batches with fewer samples per batch consistently accelerated convergence to high-fitness sequences. Incorporating protein-language-model derived zero-shot fitness priors also accelerated convergence, but only in proportion to how well each prior score correlated with the true fitness landscape. Together, these findings established generalizable design rules, favoring smaller acquisition batches and confidence-weighted priors, for engineering proteins from minimal experimental data. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=55 SRC="FIGDIR/small/744534v1_ufig1.gif" ALT="Figure 1"> View larger version (11K): org.highwire.dtl.DTLVardef@14cd238org.highwire.dtl.DTLVardef@7d8334org.highwire.dtl.DTLVardef@30e9f2org.highwire.dtl.DTLVardef@14f3a0c_HPS_FORMAT_FIGEXP M_FIG C_FIG

synthetic biology↗

TIRE-seq: an Integrated Sample Extraction and Transcriptomics Workflow for High Throughput Perturbation Studies

RNA sequencing (RNA-seq) is widely used in biomedical research, advancing our understanding of gene expression across biological systems. Traditional methods require upstream RNA extraction from biological inputs, adding time and expense to workflows. We developed TIRE-seq (Turbocapture Integrated RNA Expression Sequencing) to address these challenges. TIRE-seq integrates mRNA purification directly into library preparation, eliminating the need for a separate extraction step. This streamlined approach reduces turnaround time, minimizes sample loss, and improves data quality. A comparative study with the widely used Prime-seq protocol demonstrates TIRE-seqs superior sequencing efficiency with crude cell lysates as inputs. TIRE-seqs utility was demonstrated across three biological applications. It captured transcriptional changes in stimulated human T cells, revealing activation-associated gene expression profiles. It also identified key genes driving murine dendritic cell differentiation, providing insights into lineage commitment. Lastly, TIRE-seq analyzed the dose-response and time-course effects of temozolomide on patient-derived neurospheres, identifying differentially expressed genes and enriched pathways linked to the drugs mechanism of action. With its simplified workflow and high sequencing efficiency, TIRE-seq offers a cost-effective solution for large-scale gene expression studies across diverse biological systems.

genomics↗