bioRxiv ScienceSearch

Biology subjects

Zaitlen, N.

Publications and source records attributed to Zaitlen, N..

13 recordsLinked to original sources

A global map of RNA binding protein occupancy guides functional dissection of post-transcriptional regulation of the T cell transcriptome

RNA binding proteins (RBPs) mediate constitutive RNA metabolism and gene specific regulatory interactions. To identify RNA cis-regulatory elements, we developed GCLiPP, a biochemical technique for detecting RBP occupancy transcriptome-wide. GCLiPP sequence tags corresponded with known RBP binding sites, specifically correlating to abundant cytosolic RBPs. To demonstrate the utility of our occupancy profiles, we performed functional dissection of 3' UTRs with CRISPR/Cas9 genome editing. Two RBP occupied sites in the CD69 3' UTR destabilized the transcript of this key regulator of lymphocyte tissue egress. Comparing human Jurkat T cells and mouse primary T cells uncovered hundreds of biochemically shared peaks of GCLiPP signal across homologous regions of human and mouse 3' UTRs, including a cis-regulatory element that governs the stability of the mRNA that encodes the proto-oncogene PIM3 in both species. Our GCLiPP datasets provide a rich resource for investigation of post-transcriptional regulation in the immune system.

genomics

Reverse GWAS: Using Genetics to Identify and Model Phenotypic Subtypes

Recent and classical work has revealed biologically and medically significant subtypes in complex diseases and traits. However, relevant subtypes are often unknown, unmeasured, or actively debated, making automatic statistical approaches to subtype definition particularly valuable. We propose reverse GWAS (RGWAS) to identify and validate subtypes using genetics and multiple traits: while GWAS seeks the genetic basis of a given trait, RGWAS seeks to define trait subtypes with distinct genetic bases. Unlike existing approaches relying on off-the-shelf clustering methods, RGWAS uses a bespoke decomposition, MFMR, to model covariates, binary traits, and population structure. We use extensive simulations to show these features can be crucial for power and calibration. We validate RGWAS in practice by recovering known stress subtypes in major depressive disorder. We then show the utility of RGWAS by identifying three novel subtypes of metabolic traits. We biologically validate these metabolic subtypes with SNP-level tests and a novel polygenic test: the former recover known metabolic GxE SNPs; the latter suggests genetic heterogeneity may explain substantial missing heritability. Crucially, statins, which are widely prescribed and theorized to increase diabetes risk, have opposing effects on blood glucose across metabolic subtypes, suggesting potential have potential translational value.\n\nAuthor summaryComplex diseases depend on interactions between many known and unknown genetic and environmental factors. However, most studies aggregate these strata and test for associations on average across samples, though biological factors and medical interventions can have dramatically different effects on different people. Further, more-sophisticated models are often infeasible because relevant sources of heterogeneity are not generally known a priori. We introduce Reverse GWAS to simultaneously split samples into homogeneoues subtypes and to learn differences in genetic or treatment effects between subtypes. Unlike existing approaches to computational subtype identification using high-dimensional trait data, RGWAS accounts for covariates, binary disease traits and, especially, population structure; these features are each invaluable in extensive simulations. We validate RGWAS by recovering known genetic subtypes of major depression. We demonstrate RGWAS is practically useful in a metabolic study, finding three novel subtypes with both SNP- and polygenic-level heterogeneity. Importantly, RGWAS can uncover differential treatment response: for example, we show that statin, a common drug and potential type 2 diabetes risk factor, may have opposing subtype-specific effects on blood glucose.

genetics

Existence and implications of population variance structure

Identifying the genetic and environmental factors underlying phenotypic differences between populations is fundamental to multiple research communities. To date, studies have focused on the relationship between population and phenotypic mean. Here we consider the relationship between population and phenotypic variance, i.e., \"population variance structure.\" In addition to gene-gene and gene-environment interaction, we show that population variance structure is a direct consequence of natural selection. We develop the ancestry double generalized linear model (ADGLM), a statistical framework to jointly model population mean and variance effects. We apply ADGLM to several deeply phenotyped datasets and observe ancestry-variance associations with 12 of 44 tested traits in ~113K British individuals and 3 of 14 tested traits in ~3K Mexican, Puerto Rican, and African-American individuals. We show through extensive simulations that population variance structure can both bias and reduce the power of genetic association studies, even when principal components or linear mixed models are used. ADGLM corrects this bias and improves power relative to previous methods in both simulated and real datasets. Additionally, ADGLM identifies 17 novel genotype-variance associations across six phenotypes.

genetics

GxEMM: Extending linear mixed models to general gene-environment interactions

Gene-environment interaction (GxE) is a well-known source of non-additive inheritance. GxE can be important in applications ranging from basic functional genomics to precision medical treatment. Further, GxE effects elude inherently-linear LMMs and may explain missing heritability. We propose a simple, unifying mixed model for polygenic interactions (GxEMM) to capture the aggregate effect of small GxE effects spread across the genome. GxEMM extends existing LMMs for GxE in two important ways. First, it extends to arbitrary environmental variables, not just categorical groups. Second, GxEMM can estimate and test for environment-specific heritability. In simulations where the assumptions of existing methods do not hold, we show that GxEMM improves estimates of ordinary and GxE heritability and increases power to test for polygenic GxE. We then use GxEMM to prove that the heritability of major depression (MD) is reduced by stress, which we previously conjectured but could not prove with prior methods, and that a tail of polygenic GxE effects remains unexplained by MD GWAS.

genetics

GBAT: a gene-based association method for robust trans-gene regulation detection

Identification of trans-eQTLs has been limited by a heavy multiple testing burden, read-mapping biases, and hidden confounders. To address these issues, we developed GBAT, a powerful gene-based method that allows robust detection of trans gene regulation. Using simulated and real data, we show that GBAT drastically increases detection of trans-gene regulation over standard trans-eQTL analyses.

genetics

Genetic and environmental perturbations lead to regulatory decoherence

Correlation among traits is a fundamental feature of biological systems. From morphological characters, to transcriptional or metabolic networks, the correlations we routinely observe between traits reflect a shared regulation that remains poorly understood and difficult to study. To address this problem, we developed a new and flexible approach that allows us to identify factors associated with variation in correlation between individuals. Here, we use data from three large human cohorts to study the effects of genetic variation and environmental perturbation on correlations among mRNA transcripts and among NMR metabolites. We first show that environmental exposures (namely, infection and disease) lead to a systematic loss of correlation, which we define as decoherence. Using longitudinal data, we show that decoherent metabolites are better predictors of whether someone will develop metabolic syndrome than metabolites commonly used as biomarkers of this disease. Finally, we show that correlation itself is a trait under genetic control: specifically, we mapped and replicated hundreds of correlation QTLs, which often involve transcription factors or their known target genes. Together, this work furthers our understanding of how and why coordinated biological processes break down, and highlights the role of decoherence in disease emergence.

genomics

Tracing cellular heterogeneity in pooled genetic screens via multi-level barcoding

While pooled loss- and gain-of-function screening approaches have become increasingly popular to systematically investigate mammalian gene function, they have thus far ignored the fact that cell populations are heterogeneous. Here we introduce multi-level barcoded sgRNA libraries to (i) monitor differences in the behavior of multiplexed clonal cell lines, (ii) trace sub-clonal lineages of cells expressing the same sgRNA, (iii) derive in-sample screen replicates and (iv) reduce the number of cells and sequencing read counts required to reach statistical significance. Using our approach, we illustrate how clonal heterogeneity impairs the results of pooled genetic screens and demonstrate the ability of multi-level barcoding to resolve cellular heterogeneity related issues.

genomics

Singleton Variants Dominate the Genetic Architecture of Human Gene Expression

The vast majority of human mutations have minor allele frequencies (MAF) under 1%, with the plurality observed only once (i.e., \"singletons\"). While Mendelian diseases are predominantly caused by rare alleles, their cumulative contribution to complex phenotypes remains largely unknown. We develop and rigorously validate an approach to jointly estimate the contribution of all alleles, including singletons, to phenotypic variation. We apply our approach to transcriptional regulation, an intermediate between genetic variation and complex disease. Using whole genome DNA and lymphoblastoid cell line RNA sequencing data from 360 European individuals, we conservatively estimate that singletons contribute ~25% of cis-heritability across genes (dwarfing the contributions of other frequencies). Strikingly, the majority (~76%) of singleton heritability derives from ultra-rare variants absent from thousands of additional samples. We develop a novel inference procedure to demonstrate that our results are consistent with rampant purifying selection shaping the regulatory architecture of most human genes.

evolutionary biology

Creating ethnicity-specific reference intervals for lab tests from EHR data

The results of clinical lab tests are an essential component of medical decision-making. To guide interpretation, test results are returned with reference intervals defined by the range in which 95% of values occur in healthy individuals. Clinical laboratories often set their own reference intervals to accommodate local population and instruments variations. This approach is costly and can be biased. We describe a novel data-driven method for using electronic health record data to extract healthy patients information to define reference intervals. We found that the distributions of many clinical lab tests differ among self-identified racial and ethnic groups (SIREs) in healthy patients. Finally, we derived SIRE-specific reference intervals and provide evidence that these intervals have clinical prognostic value. Specifically, we show that for two lab tests, serum creatinine level and hemoglobin A1C, SIRE-specific reference intervals are more predictive for need for dialysis and development type 2 diabetes than existing reference intervals.\n\nOne Sentence SummaryA novel method for defining population-specific reference intervals of common clinical laboratory tests from electronical health records has better prognostic value than existing reference intervals.

bioinformatics

Adjusting For Principal Components Of Molecular Phenotypes Induces Replicating False Positives

High-throughput measurements of molecular phenotypes provide an unprecedented opportunity to model cellular processes and their impact on disease. Such highly-structured data is strongly confounded, and principal components and their variants reliably estimate latent confounders. Conditioning on PCs in downstream analyses is known to improve power and reduce multiple-testing miscalibration and is an indispensable element of thousands of published functional genomic analyses. Further clarifying this approach is of fundamental interest to the genomics and statistics communities. We uncover a novel bias induced by PC conditioning and provide an analytic, deterministic and intuitive approximation. The bias exists because PCs are, roughly, unshielded colliders on a causal path: because PCs partially incorporate a causal genotype effect on one phenotype, the genotype becomes correlated with every phenotype conditional on PCs. We empirically quantify this bias in realistic simulations. For small genetic effects, a nearly negligible bias is observed for all tested PC variants. For large genetic effects, or other differential covariates, dramatic false positives can arise. Though one PC variant (supervised SVA) largely avoids this bias, it is computationally prohibitive genome-wide; further, its immunity to this bias is novel. Our analysis informs best practices for confounder correction in genomic studies.

genetics

Decoding directional genetic dependencies through orthogonal CRISPR/Cas screens

Genetic interaction studies are a powerful approach to identify functional interactions between genes. This approach can reveal networks of regulatory hubs and connect uncharacterised genes to well-studied pathways. However, this approach has previously been limited to simple gene inactivation studies. Here, we present an orthogonal CRISPR/Cas-mediated genetic interaction approach that allows the systematic activation of one gene while simultaneously knocking out a second gene in the same cell. We have developed this concept into a quantitative and scalable combinatorial screening platform that allows the parallel interrogation of hundreds of thousands of genetic interactions. We demonstrate that the established platform works robustly to uncover genetic interactions in human cancer cells and to interpret the direction of the flow of genetic information.

molecular biology

Multiplexing droplet-based single cell RNA-sequencing using natural genetic barcodes

Droplet-based single-cell RNA-sequencing (dscRNA-seq) has enabled rapid, massively parallel profiling of transcriptomes from tens of thousands of cells. Multiplexing samples for single cell capture and library preparation in dscRNA-seq would enable cost-effective designs of differential expression and genetic studies while avoiding technical batch effects, but its implementation remains challenging. Here, we introduce an in-silico algorithm demuxlet that harnesses natural genetic variation to discover the sample identity of each cell and identify droplets containing two cells. These capabilities enable multiplexed dscRNA-seq experiments where cells from unrelated individuals are pooled and captured at higher throughput than standard workflows. To demonstrate the performance of demuxlet, we sequenced 3 pools of peripheral blood mononuclear cells (PBMCs) from 8 lupus patients. Given genotyping data for each individual, demuxlet correctly recovered the sample identity of > 99% of singlets, and identified doublets at rates consistent with previous estimates. In PBMCs, we demonstrate the utility of multiplexed dscRNA-seq in two applications: characterizing cell type specificity and inter-individual variability of cytokine response from 8 lupus patients and mapping genetic variants associated with cell type specific gene expression from 23 donors. Demuxlet is fast, accurate, scalable and could be extended to other single cell datasets that incorporate natural or synthetic DNA barcodes.

bioinformatics

Profiling adaptive immune repertoires across multiple human tissues by RNA Sequencing

Assay-based approaches provide a detailed view of the adaptive immune system by profiling T and B cell receptor repertoires. However, these methods come at a high cost and lack the scale of standard RNA sequencing (RNA-seq). Here we report the development of ImReP, a novel computational method for rapid and accurate profiling of the adaptive immune repertoire from regular RNA-Seq data. We applied it to 8,555 samples across 544 individuals from 53 tissues from the Genotype-Tissue Expression (GTEx v6) project. ImReP is able to efficiently extract TCR- and BCR- derived reads from the RNA-Seq data and accurately assemble the complementarity determining regions 3 (CDR3s), the most variable regions of B- and T-cell receptors determining their antigen specificity. Using ImReP, we have created the systematic atlas of immunological sequences for B- and T-cell repertoires across a broad range of tissue types, most of which have not been studied for B and T cell receptor repertoires. We have also examined the compositional similarities of clonal populations between the GTEx tissues to track the flow of T- and B- clonotypes across immune-related tissues, including secondary lymphoid organs and organs encompassing mucosal, exocrine, and endocrine sites. The atlas of T- and B-cell receptor receptors, freely available at https://sergheimangul.wordpress.com/atlas-immune-repertoires/, is the largest collection of CDR3 sequences and tissue types. We anticipate this recourse will enhance future studies in areas such as immunology and advance development of therapies for human diseases. ImReP is freely available at https://sergheimangul.wordpress.com/imrep/.

immunology