bioRxiv ScienceSearch

Biology subjects

Schellenberg, G. D.

Publications and source records attributed to Schellenberg, G. D..

7 recordsLinked to original sources

Inferring the molecular mechanisms of noncoding Alzheimer’s disease-associated genetic variants

Structured AbstractO_ST_ABSINTRODUCTIONC_ST_ABSWe set out to characterize the causal variants, regulatory mechanisms, tissue contexts, and target genes underlying noncoding late-onset Alzheimers Disease (LOAD)-associated genetic signals.\n\nMETHODSWe applied our INFERNO method to the IGAP genome-wide association study (GWAS) data, annotating all potentially causal variants with tissue-specific regulatory activity. Bayesian co-localization analysis of GWAS summary statistics and eQTL data was performed to identify tissue-specific target genes.\n\nRESULTSINFERNO identified enhancer dysregulation in all 19 tag regions analyzed, significant enrichments of enhancer overlaps in the immune-related blood category, and co-localized eQTL signals overlapping enhancers from the matching tissue class in ten regions (ABCA7, BIN1, CASS4, CD2AP, CD33, CELF1, CLU, EPHA1, FERMT2, ZCWPW1). We validated the allele-specific effects of several variants on enhancer function using luciferase expression assays.\n\nDISCUSSIONIntegrating functional genomics with GWAS signals yielded insights into the regulatory mechanisms, tissue contexts, and genes affected by noncoding genetic variation associated with LOAD risk.

bioinformatics

VCPA: genomic variant calling pipeline and data management tool for Alzheimer’s Disease Sequencing Project

Summary: We report VCPA, our SNP/Indel Variant Calling Pipeline and data management tool used for analysis of whole genome and exome sequencing (WGS/WES) for the Alzheimers Disease Sequencing Project. VCPA consists of two independent but linkable components: pipeline and tracking database. The pipeline is coded in Workflow Description Language and is fully optimized for the Amazon elastic compute cloud environment. This includes steps for processing raw sequence reads including read alignment, and all the way up to variant calling using GATK. The tracking database allows users to dynamically view the statuses of jobs running and the quality metrics reported by the pipeline. Users can thus monitor the production process and diagnose if any problem arises during the procedure. All quality metrics (>100 collected per processed genome) are stored in the database, thus facilitating users to compare, share and visualize the results. To summarize, VCPA is functional equivalent to the CCDG/TOPMed pipeline. Together with the dockerized database (also available as Amazon Machine Image), users can easily process any WGS/WES data on Amazon cloud with minimal installation.\n\nAvailability: VCPA is released under the MIT license and is available for academic and nonprofit use for free. The pipeline source code and step-by-step instructions are available from the National Institute on Aging Genetics of Alzheimers Disease Data Storage Site (http://www.niagads.org/VCPA).\n\nContact: yyee@pennmedicine.upenn.edu or lswang@pennmedicine.upenn.edu\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

Quality Control and Integration of Genotypes from Two Calling Pipelines for Whole Genome Sequence Data in the Alzheimer’s Disease Sequencing Project

The Alzheimers Disease Sequencing Project (ADSP) performed whole genome sequencing (WGS) of 584 subjects from 111 multiplex families at three sequencing centers. Genotype calling of single nucleotide variants (SNVs) and insertion-deletion variants (indels) was performed centrally using GATK-HaplotypeCaller and Atlas V2. The ADSP Quality Control (QC) Working Group applied QC protocols to project-level variant call format files (VCFs) from each pipeline, and developed and implemented a novel protocol, termed \"consensus calling,\" to combine genotype calls from both pipelines into a single high-quality set. QC was applied to autosomal bi-allelic SNVs and indels, and included pipeline-recommended QC filters, variant-level QC, and sample-level QC. Low-quality variants or genotypes were excluded, and sample outliers were noted. Quality was assessed by examining Mendelian inconsistencies (MIs) among 67 parent-offspring pairs, and MIs were used to establish additional genotype-specific filters for GATK calls. After QC, 578 subjects remained. Pipeline-specific QC excluded ~12.0% of GATK and 14.5% of Atlas SNVs. Between pipelines, ~91% of SNV genotypes across all QCed variants were concordant; 4.23% and 4.56% of genotypes were exclusive to Atlas or GATK, respectively; the remaining ~0.01% of discordant genotypes were excluded. For indels, variant-level QC excluded ~36.8% of GATK and 35.3% of Atlas indels. Between pipelines, ~55.6% of indel genotypes were concordant; while 10.3% and 28.3% were exclusive to Atlas or GATK, respectively; and ~0.29% of discordant genotypes were. The final WGS consensus dataset contains 27,896,774 SNVs and 3,133,926 indels and is publicly available.\n\nAbbreviationsAD, Alzheimers disease; QC, Quality Control; LSSAC, Large-Scale Sequencing and Analysis Center; Broad, Broad Institute Genomics Service; Baylor, Baylor College of Medicine Human Genome Sequencing Center; WashU, Washington University-St. Louis McDonnell Genome Institute; WGS, whole genome sequencing; WES, whole exome sequencing; indel, insertion-deletion variants; VCF, variant control format; MI, Mendelian inconsistency; MC, Mendelian consistency; GWAS, genome-wide association study; VR, referent allele read depth; DP, overall read depth; MS, mapping score; GQ, genotype quality score; Ti/Tv, Transition/Transversion; CS, concordance code

genetics

INFERNO - INFERring the molecular mechanisms of NOncoding genetic variants

The majority of variants identified by genome-wide association studies (GWAS) reside in the noncoding genome, where they affect regulatory elements including transcriptional enhancers. We propose INFERNO (INFERring the molecular mechanisms of NOncoding genetic variants), a novel method which integrates hundreds of diverse functional genomics data sources with GWAS summary statistics to identify putatively causal noncoding variants underlying association signals. INFERNO comprehensively infers the relevant tissue contexts, target genes, and downstream biological processes affected by causal variants. We apply INFERNO to schizophrenia GWAS data, recapitulating known schizophrenia-associated genes including CACNA1C and discovering novel signals related to transmembrane cellular processes.

bioinformatics

Polygenic hazard score: an enrichment marker for Alzheimer’s associated amyloid and tau deposition

BackgroundThere is an urgent need for the early identification of nondemented individuals at the highest risk of progressing to Alzheimers disease (AD) dementia for early therapeutic interventions. Our goal was to evaluate whether a recently validated polygenic hazard score (PHS) can be integrated with known in vivo CSF or PET biomarkers of amyloid or tau pathology to prospectively predict cognitive decline and clinical progression to AD dementia in nondemented older individuals.\n\nMethodsWe evaluated 347 cognitive normal (CN) and 599 mild cognitively impaired (MCI) individuals. We first investigated whether PHS can predict CSF or PET amyloid and tau deposition. We evaluated differences in positive and negative predictive values of biomarker status, as a function of PHS risk. Next, we used linear mixed-effects (LME) to examine if PHS and biomarker status in conjunction, best predict longitudinal cognitive and clinical progression. Lastly, we used survival analysis to investigate whether a combination of PHS and biomarker positivity predicts progression to AD dementia better than using PHS or biomarker positivity alone.\n\nFindingsIn CN and MCI individuals, we found that amyloid and total tau positivity systematically varies as a function of PHS. For individuals in greater than the 50th percentile PHS, the positive predictive value for amyloid approached 100%. Similarly, for individuals in less than the 25th percentile PHS, the negative predictive value for total tau approached 85%. Beyond APOE, high PHS individuals with amyloid and tau pathology showed the fastest rate of longitudinal cognitive decline and time to AD dementia progression. Among the CN subgroup, we similarly found that PHS was strongly associated with amyloid positivity and the combination of PHS and biomarker status significantly predicted longitudinal clinical progression.\n\nInterpretationAmong asymptomatic and mildly symptomatic older individuals, PHS considerably improves the predictive value of CSF or PET amyloid and tau biomarkers. Beyond APOE, PHS may be useful for risk stratification and cohort enrichment for MCI and preclinical AD therapeutic trials.

genetics

Immune-related genetic enrichment in frontotemporal dementia

BackgroundConverging evidence suggests that immune-mediated dysfunction plays an important role in the pathogenesis of frontotemporal dementia (FTD). Although genetic studies have shown that immune-associated loci are associated with increased FTD risk, a systematic investigation of genetic overlap between immune-mediated diseases and the spectrum of FTD-related disorders has not been performed.\n\nMethods and findingsUsing large genome-wide association studies (GWAS) (total n = 192,886 cases and controls) and recently developed tools to quantify genetic overlap/pleiotropy, we systematically identified single nucleotide polymorphisms (SNPs) jointly associated with FTD-related disorders namely FTD, corticobasal degeneration (CBD), progressive supranuclear palsy (PSP), and amyotrophic lateral sclerosis (ALS) - and one or more immune-mediated diseases including Crohns disease (CD), ulcerative colitis (UC), rheumatoid arthritis (RA), type 1 diabetes (T1D), celiac disease (CeD), and psoriasis (PSOR). We found up to 270-fold genetic enrichment between FTD and RA and comparable enrichment between FTD and UC, T1D, and CeD. In contrast, we found only modest genetic enrichment between any of the immune-mediated diseases and CBD, PSP or ALS. At a conjunction false discovery rate (FDR) < 0.05, we identified numerous FTD-immune pleiotropic SNPs within the human leukocyte antigen (HLA) region on chromosome 6. By leveraging the immune diseases, we also found novel FTD susceptibility loci within LRRK2 (Leucine Rich Repeat Kinase 2), TBKBP1 (TANK-binding kinase 1 Binding Protein 1), and PGBD5 (PiggyBac Transposable Element Derived 5). Functionally, we found that expression of FTD-immune pleiotropic genes (particularly within the HLA region) is altered in postmortem brain tissue from patients with frontotemporal dementia and is enriched in microglia compared to other central nervous system (CNS) cell types.\n\nConclusionsWe show considerable immune-mediated genetic enrichment specifically in FTD, particularly within the HLA region. Our genetic results suggest that for a subset of patients, immune dysfunction may contribute to risk for FTD. These findings have potential implications for clinical trials targeting immune dysfunction in patients with FTD.

genetics

Polygenic hazard scores in preclinical Alzheimer’s disease

Identifying asymptomatic older individuals at elevated risk for developing Alzheimers disease (AD) is of clinical importance. Among 1,081 asymptomatic older adults, a recently validated polygenic hazard score (PHS) significantly predicted time to AD dementia and steeper longitudinal cognitive decline, even after controlling for APOE {varepsilon}4 carrier status. Older individuals in the highest PHS percentiles showed the highest AD incidence rates. PHS predicted longitudinal clinical decline among older individuals with moderate to high CERAD (amyloid) and Braak (tau) scores at autopsy, even among APOE {varepsilon}4 non-carriers. Beyond APOE, PHS may help identify asymptomatic individuals at highest risk for developing Alzheimers neurodegeneration.

genetics