bioRxiv ScienceSearch

Biology subjects

Ertekin-Taner, N.

Publications and source records attributed to Ertekin-Taner, N..

3 recordsLinked to original sources

Large eQTL meta-analysis reveals differing patterns between cerebral cortical and cerebellar brain regions

The availability of high-quality RNA-sequencing and genotyping data of post-mortem brain collections from consortia such as CommonMind Consortium (CMC) and the Accelerating Medicines Partnership for Alzheimers Disease (AMP-AD) Consortium enable the generation of a large-scale brain cis-eQTL meta-analysis. Here we generate cerebral cortical eQTL from 1433 samples available from four cohorts (identifying >4.1 million significant eQTL for >18,000 genes), as well as cerebellar eQTL from 261 samples (identifying 874,836 significant eQTL for >10,000 genes), and provide the results as a community resource. We find substantially improved power in the meta-analysis over individual cohort analyses, particularly in comparison to the Genotype-Tissue Expression (GTEx) Project eQTL. In addition, we observed differences in eQTL patterns between cerebral and cerebellar brain regions. We provide these brain eQTL as a common resource for use across the community in research programs. As a proof of principle for their utility, we apply a colocalization analysis to identify genes underlying the GWAS association peaks for schizophrenia and identify a potentially novel gene colocalization with lncRNA RP11-677M14.2 (posterior probability of colocalization 0.975).

genetics

Systematic analysis of dark and camouflaged genes: disease-relevant genes hiding in plain sight

BackgroundThe human genome contains dark gene regions that cannot be adequately assembled or aligned using standard short-read sequencing technologies, preventing researchers from identifying mutations within these gene regions that may be relevant to human disease. Here, we identify regions that are dark by depth (few mappable reads) and others that are camouflaged (ambiguous alignment), and we assess how well long-read technologies resolve these regions. We further present an algorithm to resolve most camouflaged regions (including in short-read data) and apply it to the Alzheimers Disease Sequencing Project (ADSP; 13142 samples), as a proof of principle.\n\nResultsBased on standard whole-genome lllumina sequencing data, we identified 37873 dark regions in 5857 gene bodies (3635 protein-coding) from pathways important to human health, development, and reproduction. Of the 5857 gene bodies, 494 (8.4%) were 100% dark (142 protein-coding) and 2046 (34.9%) were [≥]5% dark (628 protein-coding). Exactly 2757 dark regions were in protein-coding exons (CDS) across 744 genes. Long-read sequencing technologies from 10x Genomics, PacBio, and Oxford Nanopore Technologies reduced dark CDS regions to approximately 45.1%, 33.3%, and 18.2% respectively. Applying our algorithm to the ADSP, we rescued 4622 exonic variants from 501 camouflaged genes, including a rare, ten-nucleotide frameshift deletion in CR1, a top Alzheimers disease gene, found in only five ADSP cases and zero controls.\n\nConclusionsWhile we could not formally assess the CR1 frameshift mutation in Alzheimers disease (insufficient sample-size), we believe it merits investigating in a larger cohort. There remain thousands of potentially important genomic regions overlooked by short-read sequencing that are largely resolved by long-read technologies.

genomics

Meta-analysis of the human brain transcriptome identifies heterogeneity across human AD coexpression modules robust to sample collection and methodological approach

Alzheimers disease (AD) is a complex and heterogenous brain disease that affects multiple inter-related biological processes. This complexity contributes, in part, to existing difficulties in the identification of successful disease-modifying therapeutic strategies. To address this, systems approaches are being used to characterize AD-related disruption in molecular state. To evaluate the consistency across these molecular models, a consensus atlas of the human brain transcriptome was developed through coexpression meta-analysis across the AMP-AD consortium. Consensus analysis was performed across five coexpression methods used to analyze RNA-seq data collected from 2114 samples across 7 brain regions and 3 research studies. From this analysis, five consensus clusters were identified that described the major sources of AD-related alterations in transcriptional state that were consistent across studies, methods, and samples. AD genetic associations, previously studied AD-related biological processes, and AD targets under active investigation were enriched in only three of these five clusters. The remaining two clusters demonstrated strong heterogeneity between males and females in AD-related expression that was consistently observed across studies. AD transcriptional modules identified by systems analysis of individual AMP-AD teams were all represented in one of these five consensus clusters except ROS/MAP-identified Module 109, which was specific for genes that showed the strongest association with changes in AD-related gene expression across consensus clusters. The other two AMP-AD transcriptional analyses reported modules that were enriched in one of the two sex-specific Consensus Clusters. The fifth cluster has not been previously identified and was enriched for genes related to proteostasis. This study provides an atlas to map across biological inquiries of AD with the goal of supporting an expansion in AD target discovery efforts.

systems biology