bioRxiv ScienceSearch

Biology subjects

Shendure, J.

Publications and source records attributed to Shendure, J..

At least 19 recordsLinked to original sources

Genome wide association with quantitative resistance phenotypes in Mycobacterium tuberculosis reveals novel resistance genes and regulatory regions

Drug resistance is threatening attempts at tuberculosis epidemic control. Molecular diagnostics for drug resistance that rely on the detection of resistance-related mutations could expedite patient care and accelerate progress in TB eradication. We performed minimum inhibitory concentration testing for 12 anti-TB drugs together with Illumina whole genome sequencing on 1452 clinical Mycobacterium tuberculosis (MTB) isolates. We then used a linear mixed model to evaluate genome wide associations between mutations in MTB genes or noncoding regions and drug resistance, followed by validation of our findings in an independent dataset of 792 patient isolates. Novel associations at 13 genomic loci were confirmed in the validation set, with 2 involving noncoding regions. We found promoter mutations to have smaller average effects on resistance levels than gene body mutations in genes where both can contribute to resistance. Enabled by a quantitative measure of resistance, we estimated the heritability of the resistance phenotype to 11 anti-TB drugs and identify a lower than expected contribution from known resistance genes. We also report the proportion of variation in resistance levels explained by the novel loci identified here. This study highlights the complexity of the genomic mechanisms associated with the MTB resistance phenotype, including the relatively large number of potentially causative or compensatory loci, and emphasizes the contribution of the noncoding portion of the genome.

evolutionary biology

Predicting mRNA abundance directly from genomic sequence using deep convolutional neural networks

Algorithms that accurately predict gene structure from primary sequence alone were transformative for annotating the human genome. Can we also predict the expression levels of genes based solely on genome sequence? Here we sought to apply deep convolutional neural networks towards this goal. Surprisingly, a model that includes only promoter sequences and features associated with mRNA stability explains 59% and 71% of variation in steady-state mRNA levels in human and mouse, respectively. This model, which we call Xpresso, more than doubles the accuracy of alternative sequence-based models, and isolates rules as predictive as models relying on ChIP-seq data. Xpresso recapitulates genome-wide patterns of transcriptional activity and predicts the influence of enhancers, heterochromatic domains, and microRNAs. Model interpretation reveals that promoter-proximal CpG dinucleotides strongly predict transcriptional activity. Looking forward, we propose the accurate prediction of cell type-specific gene expression based solely on primary sequence as a grand challenge for the field.

genomics

A combination of transcription factors mediates inducible interchromosomal pairing

Remodeling of the three-dimensional organization of a genome has been previously described (e.g. condition-specific pairing or looping), but it remains unknown which factors specify and mediate such shifts in chromosome conformation. Here we describe an assay, MAP-C (Mutation Analysis in Pools by Chromosome conformation capture), that enables the simultaneous characterization of hundreds of cis or trans-acting mutations for their effects on a chromosomal contact or loop. As a proof of concept, we applied MAP-C to systematically dissect the molecular mechanism of inducible interchromosomal pairing between HAS1pr-TDA1pr alleles in Saccharomyces yeast. We identified three transcription factors, Leu3, Sdd4 (Ypr022c), and Rgt1, whose collective binding to nearby DNA sequences is necessary and sufficient for inducible pairing between binding site clusters. Rgt1 contributes to the regulation of pairing, both through changes in expression level and through its interactions with the Tup1/Ssn6 repressor complex. HAS1pr-TDA1pr is the only locus with a cluster of binding site motifs for all three factors in both S. cerevisiae and S. uvarum genomes, but the promoter for HXT3, which contains Leu3 and Rgt1 motifs, also exhibits inducible homolog pairing. Altogether, our results demonstrate that specific combinations of transcription factors can mediate condition-specific interchromosomal contacts, and reveal a molecular mechanism for interchromosomal contacts and mitotic homolog pairing.

genomics

Functional Testing of Thousands of Osteoarthritis-Associated Variants for Regulatory Activity

To date, genome-wide association studies have implicated at least 35 loci in osteoarthritis, but due to linkage disequilibrium, we have yet to pinpoint the specific variants that underlie these associations, nor the mechanisms by which they contribute to disease risk. Here we functionally tested 1,605 single nucleotide variants associated with osteoarthritis for regulatory activity using a massively parallel reporter assay. We identified six single nucleotide polymorphisms (SNPs) with differential regulatory activity between the major and minor alleles. We show that our most significant hit, rs4730222, drives increased expression of an alternative isoform of HBP1 in a heterozygote chondrosarcoma cell line, a CRISPR-edited osteosarcoma cell line, and in chondrocytes derived from osteoarthritis patients.

genomics

crisprQTL mapping as a genome-wide association framework for cellular genetic screens

Expression quantitative trait locus (eQTL) and genome-wide association studies (GWAS) are powerful paradigms for mapping the determinants of gene expression and organismal phenotypes, respectively. However, eQTL mapping and GWAS are limited in scope (to naturally occurring, common genetic variants) and resolution (by linkage disequilibrium). Here, we present crisprQTL mapping, a framework in which large numbers of CRISPR/Cas9 perturbations are introduced to each cell on an isogenic background, followed by single-cell RNA-seq (scRNA-seq). crisprQTL mapping is analogous to conventional human eQTL studies, but with individual humans replaced by individual cells; genetic variants replaced by unique combinations of unlinked guide RNA (gRNA)-programmed perturbations per cell; and tissue-level RNA-seq of many individuals replaced by scRNA-seq of many cells. By randomly introducing gRNAs, a single population of cells can be leveraged to test for association between each perturbation and the expression of any potential target gene, analogous to how eQTL studies leverage populations of humans to test millions of genetic variants for associations with expression in a genome-wide manner. However, crisprQTL mapping is neither limited to naturally occurring, common genetic variants nor by linkage disequilibrium. As a proof-of-concept, we applied crisprQTL mapping to evaluate 1,119 candidate enhancers with no strong a priori hypothesis as to their target gene(s). Perturbations were made by a nuclease-dead Cas9 (dCas9) tethered to KRAB, and introduced at a mean allele frequency of 1.1% into a population of 47,650 profiled human K562 cells (median of 15 gRNAs identified per cell). We tested for differential expression of all genes within 1 megabase (Mb) of each candidate enhancer, effectively evaluating 17,584 potential enhancer-target gene relationships within a single experiment. At an empirical false discovery rate (FDR) of 10%, we identify 128 cis crisprQTLs (11%) whose targeting resulted in downregulation of 105 nearby genes. crisprQTLs were strongly enriched for proximity to their target genes (median 34.3 kilobases (Kb)) and the strength of H3K27ac, p300, and lineage-specific transcription factor (TF) ChIP-seq peaks. Our results establish the power of the eQTL mapping paradigm as applied to programmed variation in populations of cells, rather than natural variation in populations of individuals. We anticipate that crisprQTL mapping will facilitate the comprehensive elucidation of the cis-regulatory architecture of the human genome.

genomics

Accurate functional classification of thousands of BRCA1 variants with saturation genome editing

Variants of uncertain significance (VUS) fundamentally limit the utility of genetic information in a clinical setting. The challenge of VUS is epitomized by BRCA1, a tumor suppressor gene integral to DNA repair and genomic stability. Germline BRCA1 loss-of-function (LOF) variants predispose women to early-onset breast and ovarian cancers. Although BRCA1 has been sequenced in millions of women, the risk associated with most newly observed variants cannot be definitively assigned. Data sharing attenuates this problem but it is unlikely to solve it, as most newly observed variants are exceedingly rare. In lieu of genetic evidence, experimental approaches can be used to functionally characterize VUS. However, to date, functional studies of BRCA1 VUS have been conducted in a post hoc, piecemeal fashion. Here we employ saturation genome editing to assay 96.5% of all possible single nucleotide variants (SNVs) in 13 exons that encode functionally critical domains of BRCA1. Our assay measures cellular fitness in a haploid human cell line whose survival is dependent on intact BRCA1 function. The resulting function scores for nearly 4,000 SNVs are bimodally distributed and almost perfectly concordant with established assessments of pathogenicity. Sequence-function maps enhanced by parallel measurements of variant effects on mRNA levels reveal mechanisms by which loss-of-function SNVs arise. Hundreds of missense SNVs critical for protein function are identified, as well as dozens of exonic and intronic SNVs that compromise BRCA1 function by disrupting splicing or transcript stability. We predict that these function scores will be directly useful for the clinical interpretation of cancer risk based on BRCA1 sequencing. Furthermore, we propose that this paradigm can be extended to overcome the challenge of VUS in other genes in which genetic variation is clinically actionable.

genomics

A multiplexed homology-directed DNA repair assay reveals the impact of ~1,700 BRCA1 variants on protein function

Loss-of-function mutations in BRCA1 confer a predisposition to breast and ovarian cancer. Genetic testing for mutations in the BRCA1 gene frequently reveals a missense variant for which the impact on the molecular function of the BRCA1 protein is unknown. Functional BRCA1 is required for homology directed repair (HDR) of double-strand DNA breaks, a key activity for maintaining genome integrity and tumor suppression. Here we describe a multiplex HDR reporter assay to simultaneously measure the effect of hundreds of variants of BRCA1 on its role in DNA repair. Using this assay, we measured the effects of ~1,700 amino acid substitutions in the first 302 residues of BRCA1. Benchmarking these results against variants with known effects, we demonstrate accurate discrimination of loss-of-function versus benign variants. We anticipate that this assay can be used to functionally characterize BRCA1 missense variants at scale, even before the variants are observed in results from genetic testing.

genomics

Functional Characterization of Enhancer Evolution in the Primate Lineage

BackgroundEnhancers play an important role in morphological evolution and speciation by controlling the spatiotemporal expression of genes. Due to technological limitations, previous efforts to understand the evolution of enhancers in primates have typically studied many enhancers at low resolution, or single enhancers at high resolution. Although comparative genomic studies reveal large-scale turnover of enhancers, a specific understanding of the molecular steps by which mammalian or primate enhancers evolve remains elusive.\n\nResultsWe identified candidate hominoid-specific liver enhancers from H3K27ac ChIP-seq data. After locating orthologs in 11 primates spanning [~]40 million years, we synthesized all orthologs as well as computational reconstructions of 9 ancestral sequences for 348 \"active tiles\" of 233 putative enhancers. We concurrently tested all sequences (20 per tile) for regulatory activity with STARR-seq in HepG2 cells, with the goal of characterizing the evolutionary-functional trajectories of each enhancer. We observe groups of enhancer tiles with coherent trajectories, most of which can be explained by one or two mutational events per tile. We quantify the correlation between the number of mutations along a branch and the magnitude of change in functional activity. Finally, we identify 57 mutations that correlate with functional changes; these are enriched for cytosine deamination events within CpGs, compared to background events.\n\nConclusionsWe characterized the evolutionary-functional trajectories of hundreds of liver enhancers throughout the primate phylogeny. We observe subsets of regulatory sequences that appear to have gained or lost activity at various positions in the primate phylogeny. We use these data to quantify the relationship between sequence and functional divergence, and to identify CpG deamination as a potentially important force in driving changes in enhancer activity during primate evolution.

genomics

On the design of CRISPR-based single cell molecular screens

Several groups recently reported coupling CRISPR/Cas9 perturbations and single cell RNA-seq as a potentially powerful approach for forward genetics. Here we demonstrate that vector designs for such screens that rely on cis linkage of guides and distally located barcodes suffer from swapping of intended guide-barcode associations at rates approaching 50% due to template switching during lentivirus production, greatly reducing sensitivity. We optimize a published strategy, CROP-seq, that instead uses a Pol II transcribed copy of the sgRNA sequence itself, doubling the rate at which guides are assigned to cells to 94%. We confirm this strategy performs robustly and further explore experimental best practices for CRISPR/Cas9-based single cell molecular screens.

genomics

Massively parallel dissection of human accelerated regions in human and chimpanzee neural progenitors

Using machine learning (ML), we interrogated the function of all human-chimpanzee variants in 2,645 Human Accelerated Regions (HARs), some of the fastest evolving regions of the human genome. We predicted that 43% of HARs have variants with large opposing effects on chromatin state and 14% on neurodevelopmental enhancer activity. This pattern, consistent with compensatory evolution, was confirmed using massively parallel reporter assays in human and chimpanzee neural progenitor cells. The species-specific enhancer activity of assayed HARs was accurately predicted from the presence and absence of transcription factor footprints in each species. Despite these striking cis effects, activity of a given HAR sequence was nearly identical in human and chimpanzee cells. These findings suggest that HARs did not evolve to compensate for changes in the trans environment but instead altered their ability to bind factors present in both species. Thus, ML prioritized variants with functional effects on human neurodevelopment and revealed an unexpected reason why HARs may have evolved so rapidly.

evolutionary biology

Multiplex Assessment of Protein Variant Abundance by Massively Parallel Sequencing

Determining the pathogenicity of human genetic variants is a critical challenge, and functional assessment is often the only option. Experimentally characterizing millions of possible missense variants in thousands of clinically important genes will likely require generalizable, scalable assays. Here we describe Variant Abundance by Massively Parallel Sequencing (VAMP-seq), which measures the effects of thousands of missense variants of a protein on intracellular abundance in a single experiment. We apply VAMP-seq to quantify the abundance of 7,595 single amino acid variants of two proteins, PTEN and TPMT, in which functional variants are clinically actionable. We identify 1,079 PTEN and 805 TPMT single amino acid variants that result in low protein abundance, and may be pathogenic or alter drug metabolism, respectively. We observe selection for low-abundance PTEN variants in cancer, and our abundance data suggest that a PTEN variant accounting for ~10% of PTEN missense variants in melanomas functions via a dominant negative mechanism. Finally, we demonstrate that VAMP-seq can be applied to other genes, highlighting its potential as a generalizable assay for characterizing missense variants.

genetics

Dynamic reorganization of nuclear architecture during human cardiogenesis

While chromosomal architecture varies among cell types, little is known about how this organization is established or its role in development. We integrated Hi-C, RNA-seq and ATAC-seq during cardiac differentiation from human pluripotent stem cells to generate a comprehensive profile of chromosomal architecture. We identified active and repressive domains that are dynamic during cardiogenesis and recapitulate in vivo cardiomyocytes. During differentiation, heterochromatic regions condense in cis. In contrast, many cardiac-specific genes, such as TTN (titin), decompact and transition to an active compartment coincident with upregulation. Moreover, we identify a network of genes, including TTN, that share the heart-specific splicing factor, RBM20, and become associated in trans during differentiation, suggesting the existence of a 3D nuclear splicing factory. Our results demonstrate both the dynamic nature in nuclear architecture and provide insights into how developmental genes are coordinately regulated.\n\nOne Sentence SummaryThe three-dimensional structure of the human genome is dynamically regulated both globally and locally during cardiogenesis.

genomics

Simultaneous single-cell profiling of lineages and cell types in the vertebrate brain by scGESTALT

Hundreds of cell types are generated during development, but their lineage relationships are largely elusive. Here we report a technology, scGESTALT, which combines cell type identification by single-cell RNA sequencing with lineage recording by cumulative barcode editing. We sequenced ~60,000 transcriptomes from the juvenile zebrafish brain and identified more than 100 cell types and marker genes. We engineered an inducible system that combines early and late barcode editing and isolated thousands of single-cell transcriptomes and their associated barcodes. The large diversity of edited barcodes and cell types enabled the generation of lineage trees with hundreds of branches. Inspection of lineage trajectories identified restrictions at the level of cell types and brain regions and helped uncover gene expression cascades during differentiation. These results establish scGESTALT as a new and widely applicable tool to simultaneously characterize the molecular identities and lineage histories of thousands of cells during development and disease.

developmental biology

FlashFry: a fast and flexible tool for large-scale CRISPR target design

FlashFry is a fast and flexible command-line tool for characterizing large numbers of CRISPR target sequences. While several CRISPR web application exist, genome-wide knockout studies, noncoding deletion scans, and other large-scale studies or methods development projects require a simple and lightweight framework that can quickly discover and score thousands of candidates guides targeting an arbitrary DNA sequence. With FlashFry, users can specify an unconstrained number of mismatches to putative off-targets, richly annotate discovered sites, and tag potential guides with commonly used on target and off-target scoring metrics. FlashFry runs at speeds comparable to widely used genome-wide sequence aligners, and output is provided as an easy-to-manipulate text file.\n\nAvailabilityFlashFry is written in Scala and bundled as a stand-alone Jar file, easily run on any system with an installed Java virtual machine (JVM). The tool is freely licensed under version 3 of the GPL, and code, documentation, and tutorials are available on the GitHub page: http://aaronmck.github.io/FlashFry/

bioinformatics

Using DNase Hi-C techniques to map global and local three-dimensional genome architecture at high resolution

The folding and three-dimensional (3D) organization of chromatin in the nucleus critically impacts genome function. The past decade has witnessed rapid advances in genomic tools for delineating 3D genome architecture. Among them, chromosome conformation capture (3C)-based methods such as Hi-C are the most widely used techniques for mapping chromatin interactions. However, traditional Hi-C protocols rely on restriction enzymes (REs) to fragment chromatin and are therefore limited in resolution. We recently developed DNase Hi-C for mapping 3D genome organization, which uses DNase I for chromatin fragmentation. DNase Hi-C overcomes RE-related limitations associated with traditional Hi-C methods, leading to improved methodological resolution. Furthermore, combining this method with DNA capture technology provides a high-throughput approach (targeted DNase Hi-C) that allows for mapping fine-scale chromatin architecture at exceptionally high resolution. Hence, targeted DNase Hi-C will be valuable for delineating the physical landscapes of cis-regulatory networks that control gene expression and for characterizing phenotype-associated chromatin 3D signatures. Here, we provide a detailed description of method design and step-by-step working protocols for these two methods.\n\nHighlightsO_LIDNase Hi-C, a method for comprehensive mapping of chromatin contacts on a whole-genome scale, is based on random chromatin fragmentation by DNase I digestion instead of sequence-specific restriction enzyme (RE) digestion.\nC_LIO_LITargeted DNase Hi-C, which combines DNase Hi-C with DNA capture technology, is a high-throughput method for mapping fine-scale chromatin architecture of genomic loci of interest at a resolution comparable to that of genomic annotations of functional elements.\nC_LIO_LIDNase Hi-C and targeted DNase Hi-C provide the first high-throughput way to overcome the RE-digestion-associated resolution limit of 3C-based methods.\nC_LIO_LIStep-by-step whole-genome and targeted DNase Hi-C protocols for mapping global and local 3D genome architecture, respectively, are described.\nC_LI

genomics

The cis-regulatory dynamics of embryonic development at single cell resolution

Single cell measurements of gene expression are providing new insights into lineage commitment, yet the regulatory changes underlying individual cell trajectories remain elusive. Here, we profiled chromatin accessibility in over 20,000 single nuclei across multiple stages of Drosophila embryogenesis. Our data reveal heterogeneity in the regulatory landscape prior to gastrulation that reflects anatomical position, a feature that aligns with future cell fate. During mid embryogenesis, tissue granularity emerges such that cell types can be inferred by their chromatin accessibility, while maintaining a signature of their germ layer of origin. We identify over 30,000 distal elements with tissue-specific accessibility. Using transgenic embryos, we tested the germ layer specificity of a subset of predicted enhancers, achieving near-perfect accuracy. Overall, these data demonstrate the power of shotgun single cell profiling of embryos to resolve dynamic changes in open chromatin during development, and to uncover the cis-regulatory programs of germ layers and cell types.

genomics

Scalable and efficient single-cell DNA methylation sequencing by combinatorial indexing.

Here we present a novel method: single-cell combinatorial indexing for methylation analysis (sci-MET), which is the first highly scalable assay for whole genome methylation profiling of single cells. We use sci-MET to produce 2,697 total single-cell bisulfite sequencing libraries and achieve read alignment rates of 69 {+/-} 7%, comparable to those of bulk cell methods. As a proof of concept, we applied sci-MET to successfully deconvolve the cellular identity of a mixture of three human cell lines.

genomics

Chromatin accessibility dynamics of myogenesis at single cell resolution

Over a million DNA regulatory elements have been cataloged in the human genome, but linking these elements to the genes that they regulate remains challenging. We introduce Cicero, a statistical method that connects regulatory elements to target genes using single cell chromatin accessibility data. We apply Cicero to investigate how thousands of dynamically accessible elements orchestrate gene regulation in differentiating myoblasts. Groups of co-accessible regulatory elements linked by Cicero meet criteria of \"chromatin hubs\", in that they are physically proximal, interact with a common set of transcription factors, and undergo coordinated changes in histone marks that are predictive of gene expression. Pseudotemporal analysis revealed a subset of elements bound by MYOD in myoblasts that exhibit early opening, potentially serving as the initial sites of recruitment of chromatin remodeling and histone-modifying enzymes. The methodological framework described here constitutes a powerful new approach for elucidating the architecture, grammar and mechanisms of cis-regulation on a genome-wide basis.

genomics