bioRxiv Science⌕ Search

Biology subjects

Yan, R. E.

Publications and source records attributed to Yan, R. E..

4 recordsLinked to original sources

Rational design of synthetic proteins using a genome-scale CRISPR screen

Protein structure prediction using deep learning has revolutionized protein design. Yet, our understanding of protein function remains a key limitation for designing novel proteins that perform complex biological tasks. Here, we adopt a massively-parallel, function-first approach to rationally design synthetic proteins. Using genome-scale CRISPR activation, we overexpress [~]19,000 human proteins and measure their impact on precise gene editing. We identify over 800 native proteins that promote homology-directed repair. Using top candidates, we then design synthetic genome editors -- Targeted Repair fUsion Editors (TruEditors) -- by fusing full-length proteins or smaller core domains to the Cas9 nuclease. We develop 12 unique TruEditors that improve precise gene editing in diverse cell types and at genomic loci where existing methods for precise gene editing fail. Using affinity proteomics, we show that these synthetic proteins work by coordinating with endogenous DNA repair complexes. The delivery of TruEditors via mRNA more than doubles the rate of chimeric antigen receptor (CAR) insertion into the TRAC locus of primary human T cells, enhancing CAR T cell-directed tumor cell killing, and improves precise editing in human pluripotent stem cells more than three-fold. Overall, our study demonstrates that genome-wide protein overexpression screens can guide the rational design of synthetic proteins for specific biological tasks.

bioengineering↗

Transcriptome-wide profiling of alternative splicing regulators with CRISPore-seq

Alternative splicing creates diverse RNA isoforms from individual genes, yet single-cell CRISPR screens are limited to gene-level quantification and cannot detect changes in alternative splicing and transcript isoforms. To overcome this limitation, we develop CRISPore-seq, which couples massively-parallel CRISPR perturbations with joint short- and long-read transcriptomics. CRISPore-seq simultaneously captures genetic perturbations and expression of genes, full-length transcripts and surface proteins in single cells. CRISPore-seq long reads identify 80% more transcript isoforms than short reads. Nearly all long reads map to unique transcript isoforms -- in contrast to existing single-cell perturbation methods, which rarely distinguish specific isoforms. Using CRISPore-seq, we knock-down 15 different RNA-binding proteins (RBPs) and identify thousands of perturbation-driven alternative splicing events (ASEs). We find that exon skipping is the most common ASE observed and that skipped exons are enriched for binding sites of perturbed RBPs. Loss of the Nager syndrome-associated spliceosomal factor SF3B4 triggers skipping of exon 2 in the cell-cycle regulator CCND1, preventing formation of a complex with CDK6 and blocking the G1-S transition. After rescue with a CCND1 isoform containing the skipped exon, both holoenzyme complex formation and cell proliferation are restored. By linking genes to transcriptional phenotypes with isoform-level resolution, CRISPore-seq is a highly scalable tool for understanding the impact of genetic perturbations on the human transcriptome.

genomics↗

Machine learning-predicted chromatin organization landscape across pediatric tumors

Structural variants (SVs) are increasingly recognized as important contributors to oncogenesis through their effects on 3D genome folding. Recent advances in whole-genome sequencing have enabled large-scale profiling of SVs across diverse tumors, yet experimental characterization of their individual impact on genome folding remains infeasible. Here, we leveraged a convolutional neural network, Akita, to predict disruptions in genome folding caused by somatic SVs identified in 61 tumor types from the Childrens Brain Tumor Network dataset. Our analysis reveals significant variability in SV-induced disruptions across tumor types, with the most disruptive SVs coming from lymphomas and sarcomas, metastatic tumors, and germline cell tumors. Dimensionality reduction of disruption scores identified five recurrently disrupted regions enriched for high-impact SVs across multiple tumors. Some of these regions are highly disrupted despite not being highly mutated, and harbor tumor-associated genes and transcriptional regulators. To further interpret the functional relevance of high-scoring SVs, we integrated epigenetic data and developed a modified Activity-by-Contact scoring approach to prioritize SVs with disrupted genome contacts at active enhancers. This method highlighted highly disruptive SVs near key oncogenes, as well as novel candidate loci potentially implicated in tumorigenesis. These findings highlight the utility of machine learning for identifying novel SVs, loci, and genetic mechanisms contributing to pediatric cancers. This framework provides a foundation for future studies linking SV-driven regulatory changes to cancer pathogenesis.

genomics↗

Paired CRISPR screens to map gene regulation in cis and trans

Recent massively-parallel approaches to decipher gene regulatory circuits have focused on the discovery of either cis-regulatory elements (CREs) or trans-acting factors. Here, we develop a scalable approach that pairs cis- and trans-regulatory CRISPR screens to systematically dissect how the key immune checkpoint PD-L1 is regulated. In human pancreatic ductal adenocarcinoma (PDAC) cells, we tile the PD-L1 locus using [~]25,000 CRISPR perturbations in constitutive and IFN{gamma}-stimulated conditions. We discover 67 enhancer- or repressor-like CREs and show that distal CREs tend to contact the promoter of PD-L1 and related genes. Next, we measure how loss of all [~]2,000 transcription factors (TFs) in the human genome impacts PD-L1 expression and, using this, we link specific TFs to individual CREs and reveal novel PD-L1 regulatory circuits. For one of these regulatory circuits, we confirm the binding of predicted trans-factors (SRF and BPTF) using CUT&RUN and show that loss of either the CRE or TFs potentiates the anti-cancer activity of primary T cells engineered with a chimeric antigen receptor. Finally, we show that expression of these TFs correlates with PD-L1 expression in vivo in primary PDAC tumors and that somatic mutations in TFs can alter response and overall survival in immune checkpoint blockade-treated patients. Taken together, our approach establishes a generalizable toolkit for decoding the regulatory landscape of any gene or locus in the human genome, yielding insights into gene regulation and clinical impact.

genomics↗