bioRxiv ScienceSearch

Biology subjects

Bray, N. L.

Publications and source records attributed to Bray, N. L..

7 recordsLinked to original sources

Expression reflects population structure

Population structure in genotype data has been extensively studied, and is revealed by looking at the principal components of the genotype matrix. However, no similar analysis of population structure in gene expression data has been conducted, in part because a naive principal components analysis of the gene expression matrix does not cluster by population. We identify a linear projection that reveals population structure in gene expression data. Our approach relies on the coupling of the principal components of genotype to the principal components of gene expression via canonical correlation analysis. Futhermore, we analyze the variance of each gene within the projection matrix to determine which genes significantly influence the projection. We identify thousands of significant genes, and show that a number of the top genes have been implicated in diseases that disproportionately impact African Americans.\n\nAuthor SummaryHigh dimensional, multi-modal genomics datasets are becoming increasingly common, which warrants investigation into analysis techniques that can reveal structure in the data without over-fitting. Here, we show that the coupling of principal component analysis to canonical correlation analysis offers an efficient approach to exploratory analysis of this kind of data. We apply this method to the GEUVADIS dataset of genotype and gene expression values of European and Yoruban individuals, finding as-of-yet unstudied population structure in the gene expression values. Moreover, many of the top genes identified by our method have been previously implicated in diseases that disproportionately impact African Americans.

genomics

Controlled cycling and quiescence enables homology directed repair in engraftment-enriched adult hematopoietic stem and progenitor cells

Hematopoietic stem cells (HSCs) are the source of all blood components, and genetic defects in these cells are causative of disorders ranging from severe combined immunodeficiency to sickle cell disease. However, genome editing of long-term repopulating HSCs to correct mutated alleles has been challenging. HSCs have the ability to either be quiescent or cycle, with the former linked to stemness and the latter involved in differentiation. Here we investigate the link between cell cycle status and genome editing outcomes at the causative codon for sickle cell disease in adult human CD34+ hematopoietic stem and progenitor cells (HSPCs). We show that quiescent HSPCs that are immunophenotypically enriched for engrafting stem cells predominantly repair Cas9-induced double strand breaks (DSBs) through an error-prone non-homologous end-joining (NHEJ) pathway and exhibit almost no homology directed repair (HDR). By contrast, non-quiescent cycling stem-enriched cells repair Cas9 DSBs through both error-prone NHEJ and fidelitous HDR. Pre-treating bulk CD34+ HSPCs with a combination of mTOR and GSK-3 inhibitors to induce quiescence results in complete loss of HDR in all cell subtypes. We used these compounds, which were initially developed to maintain HSCs in culture, to create a new strategy for editing adult human HSCs. CD34+ HSPCs are edited, allowed to briefly cycle to accumulate HDR alleles, and then placed back in quiescence to maintain stemness, resulting in 6-fold increase in HDR/NHEJ ratio in quiescent, stem-enriched cells. Our results reveal the fundamental tension between quiescence and editing in human HSPCs and suggests strategies to manipulate HSCs during therapeutic genome editing.

cell biology

Gene-level differential analysis at transcript-level resolution

Gene-level differential expression analysis based on RNA-Seq is more robust, powerful and biologically actionable than transcript-level differential analysis. However aggregation of transcript counts prior to analysis results can mask transcript-level dynamics. We demonstrate that aggregating the results of transcript-level analysis allow for gene-level analysis with transcript-level resolution. We also show that p-value aggregation methods, typically used for meta-analyses, greatly increase the sensitivity of gene-level differential analyses. Furthermore, such aggregation can be applied directly to transcript compatibility counts obtained during pseudoalignment, thereby allowing for rapid and accurate model-free differential testing. The methods are general, allowing for testing not only of genes but also of any groups of transcripts, and we showcase an example where we apply them to perturbation analysis of gene ontologies.

bioinformatics

Fusion detection and quantification by pseudoalignment

RNA sequencing in cancer cells is a powerful technique to detect chromosomal rearrangements, allowing for de novo discovery of actively expressed fusion genes. Here we focus on the problem of detecting gene fusions from raw sequencing data, assembling the reads to define fusion transcripts and their associated breakpoints, and quantifying their abundances. Building on the pseudoalignment idea that simplifies and accelerates transcript quantification, we introduce a novel approach to fusion detection based on inspecting paired reads that cannot be pseudoaligned due to conflicting matches. The method and software, called pizzly, filters false positives, assembles new transcripts from the fusion reads, and reports candidate fusions. With pizzly, fusion detection from raw RNA-Seq reads can be performed in a matter of minutes, making the program suitable for the analysis of large cancer gene expression databases and for clinical use. pizzly is available at https://github.com/pmelsted/pizzly

bioinformatics

CRISPR-Cas9 Genome Editing In Human Cells Works Via The Fanconi Anemia Pathway

CRISPR-Cas9 genome editing creates targeted double strand breaks (DSBs) in eukaryotic cells that are processed by cellular DNA repair pathways. Co-administration of single stranded oligonucleotide donor DNA (ssODN) during editing can result in high-efficiency (>20%) incorporation of ssODN sequences into the break site. This process is commonly referred to as homology directed repair (HDR) and here referred to as single stranded template repair (SSTR) to distinguish it from repair using a double stranded DNA donor (dsDonor). The high efficacy of SSTR makes it a promising avenue for the treatment of genetic diseases1,2, but the genetic basis of SSTR editing is still unclear, leaving its use a mostly empiric process. To determine the pathways underlying SSTR in human cells, we developed a coupled knockdown-editing screening system capable of interrogating multiple editing outcomes in the context of thousands of individual gene knockdowns. Unexpectedly, we found that SSTR requires multiple components of the Fanconi Anemia (FA) repair pathway, but does not require Rad51-mediated homologous recombination, distinguishing SSTR from repair using dsDonors. Knockdown of FA genes impacts SSTR without altering break repair by non-homologous end joining (NHEJ) in multiple human cell lines and in neonatal dermal fibroblasts. Our results establish an unanticipated and central role for the FA pathway in templated repair from single stranded DNA by human cells. Therapeutic genome editing has been proposed to treat genetic disorders caused by deficiencies in DNA repair, including Fanconi Anemia. Our data imply that patient genotype and/or transcriptome profoundly impact the effectiveness of gene editing treatments and that adjuvant treatments to bias cells towards FA repair pathways could have considerable therapeutic value.

cell biology

Disabling Cas9 by an anti-CRISPR DNA mimic

CRISPR-Cas9 gene editing technology is derived from a microbial adaptive immune system, where bacteriophages are often the intended target. Natural inhibitors of CRISPR-Cas9 enable phages to evade immunity and show promise in controlling Cas9-mediated gene editing in human cells. However, the mechanism of CRISPR-Cas9 inhibition is not known and the potential applications for Cas9 inhibitor proteins in mammalian cells has not fully been established. We show here that the anti-CRISPR protein AcrIIA4 binds only to assembled Cas9-single guide RNA (sgRNA) complexes and not to Cas9 protein alone. A 3.9 [A] resolution cryo-EM structure of the Cas9-sgRNA-AcrIIA4 complex revealed that the surface of AcrIIA4 is highly acidic and binds with 1:1 stoichiometry to a region of Cas9 that normally engages the DNA protospacer adjacent motif (PAM). Consistent with this binding mode, order-of-addition experiments showed that AcrIIA4 interferes with DNA recognition but has no effect on pre-formed Cas9-sgRNA-DNA complexes. Timed delivery of AcrIIA4 into human cells as either protein or expression plasmid allows on-target Cas9-mediated gene editing while reducing off-target edits. These results provide a mechanistic understanding of AcrIIA4 function and demonstrate that inhibitors can modulate the extent and outcomes of Cas9-mediated gene editing.

biochemistry

Discovery of an autoimmunity-associated IL2RA enhancer by unbiased targeting of transcriptional activation

The majority of genetic variants associated with common human diseases map to enhancers, non-coding elements that shape cell type-specific transcriptional programs and responses to specific extracellular cues 1-3. In order to understand the mechanisms by which non-coding genetic variation contributes to disease, systematic mapping of functional enhancers and their biological contexts is required. Here, we develop an unbiased discovery platform that can identify enhancers for a target gene without prior knowledge of their native functional context. We used tiled CRISPR activation (CRISPRa) to synthetically recruit transcription factors to sites across large genomic regions (>100 kilobases) surrounding two key autoimmunity risk loci, CD69 and IL2RA (interleukin-2 receptor alpha; CD25). We identified several CRISPRa responsive elements (CaREs) with stimulation-dependent enhancer activity, including an IL2RA enhancer that harbors an autoimmunity risk variant. Using engineered mouse models and genome editing of human primary T cells, we found that sequence perturbation of the disease-associated IL2RA enhancer does not block IL2RA expression, but rather delays the timing of gene activation in response to specific extracellular signals. This work develops an approach to rapidly identify functional enhancers within non-coding regions, decodes a key human autoimmunity association, and suggests a general mechanism by which genetic variation can cause immune dysfunction.

genetics