bioRxiv Science⌕ Search

Biology subjects

Boger, R. S.

Publications and source records attributed to Boger, R. S..

3 recordsLinked to original sources

Rapid two-step target capture ensures efficient CRISPR-Cas9-guided genome editing

RNA-guided CRISPR-Cas enzymes initiate programmable genome editing by recognizing a 20-base-pair DNA sequence adjacent to a short protospacer-adjacent motif (PAM). To uncover the molecular determinants of high-efficiency editing, we conducted biochemical, biophysical and cell-based assays on S. pyogenes Cas9 (SpyCas9) variants with wide-ranging genome editing efficiencies that differ in PAM binding specificity. Our results show that reduced PAM specificity causes persistent non-selective DNA binding and recurrent failures to engage the target sequence through stable guide RNA hybridization, leading to reduced genome editing efficiency in cells. These findings reveal a fundamental trade-off between broad PAM recognition and genome editing effectiveness. We propose that high-efficiency RNA-guided genome editing relies on an optimized two-step target capture process, where selective but low-affinity PAM binding precedes rapid DNA unwinding. This model provides a foundation for engineering more effective CRISPR-Cas and related RNA-guided genome editors.

biophysics↗

Functional protein mining with conformal guarantees

1Molecular structure prediction and homology detection provide a promising path to discovering new protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a novel approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of new proteins with likely desirable functional properties.

bioinformatics↗

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. While advances in deep learning enable the prediction of accurate protein structural models, RNA structure prediction is not possible at present due to a lack of abundant high-quality reference data1. Furthermore, available sequence data are generally not associated with organismal phenotypes that could inform RNA function2-4. We created GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB)5. GARNET links RNA sequences derived from GTDB genomes to experimental and predicted optimal growth temperatures of GTDB reference organisms. This enables construction of deep and diverse RNA sequence alignments to be used for machine learning. Using GARNET, we define the minimal requirements for a sequence- and structure-aware RNA generative model. We also develop a GPT-like language model for RNA in which overlapping triplet tokenization provides optimal encoding. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identified mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB- derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

synthetic biology↗