bioRxiv ScienceSearch

Biology subjects

Kramer, R.

Publications and source records attributed to Kramer, R..

3 recordsLinked to original sources

BDQC: a general-purpose analytics tool for domain-blind validation of Big Data

Translational biomedical research is generating exponentially more data: thousands of whole-genome sequences (WGS) are now available; brain data are doubling every two years. Analyses of Big Data, including imaging, genomic, phenotypic, and clinical data, present qualitatively new challenges as well as opportunities. Among the challenges is a proliferation in ways analyses can fail, due largely to the increasing length and complexity of processing pipelines. Anomalies in input data, runtime resource exhaustion or node failure in a distributed computation can all cause pipeline hiccups that are not necessarily obvious in the output. Flaws that can taint results may persist undetected in complex pipelines, a danger amplified by the fact that research is often concurrent with the development of the software on which it depends. On the positive side, the huge sample sizes increase statistical power, which in turn can shed new insight and motivate innovative analytic approaches. We have developed a framework for Big Data Quality Control (BDQC) including an extensible set of heuristic and statistical analyses that identify deviations in data without regard to its meaning (domain-blind analyses). BDQC takes advantage of large sample sizes to classify the samples, estimate distributions and identify outliers. Such outliers may be symptoms of technology failure (e.g., truncated output of one step of a pipeline for a single genome) or may reveal unsuspected \" signal\" in the data (e.g., evidence of aneuploidy in a genome). We have applied the framework to validate real-world WGS analysis pipelines. BDQC successfully identified data outliers representing various failure classes, including genome analyses missing a whole chromosome or part thereof, hidden among thousands of intermediary output files. These failures could then be resolved by reanalyzing the affected samples. BDQC both identified hidden flaws as well as yielded new insights into the data. BDQC is designed to complement quality software development practices. There are multiple benefits from the application of BDQC at all pipeline stages. By verifying input correctness, it can help avoid expensive computations on flawed data. Analysis of intermediary and final results facilitates recovery from aberrant termination of processes. All these computationally inexpensive verifications reduce cryptic analytical artifacts that could seriously preclude clinical-grade genome interpretation. BDQC is available at https://github.com/ini-bdds/bdqc.

bioinformatics

Gene expression imputation across multiple brain regions reveals schizophrenia risk throughout development.

Transcriptomic imputation approaches offer an opportunity to test associations between disease and gene expression in otherwise inaccessible tissues, such as brain, by combining eQTL reference panels with large-scale genotype data. These genic associations could elucidate signals in complex GWAS loci and may disentangle the role of different tissues in disease development. Here, we use the largest eQTL reference panel for the dorso-lateral pre-frontal cortex (DLPFC), collected by the CommonMind Consortium, to create a set of gene expression predictors and demonstrate their utility. We applied these predictors to 40,299 schizophrenia cases and 65,264 matched controls, constituting the largest transcriptomic imputation study of schizophrenia to date. We also computed predicted gene expression levels for 12 additional brain regions, using publicly available predictor models from GTEx. We identified 413 genic associations across 13 brain regions. Stepwise conditioning across the genes and tissues identified 71 associated genes (67 outside the MHC), with the majority of associations found in the DLPFC, and of which 14/67 genes did not fall within previously genome-wide significant loci. We identified 36 significantly enriched pathways, including hexosaminidase-A deficiency, and multiple pathways associated with porphyric disorders. We investigated developmental expression patterns for all 67 non-MHC associated genes using BRAINSPAN, and identified groups of genes expressed specifically pre-natally or post-natally.

genetics

Transcriptomic Imputation of Bipolar Disorder and Bipolar subtypes reveals 29 novel associated genes

Bipolar disorder is a complex neuropsychiatric disorder presenting with episodic mood disturbances. In this study we use a transcriptomic imputation approach to identify novel genes and pathways associated with bipolar disorder, as well as three diagnostically and genetically distinct subtypes. Transcriptomic imputation approaches leverage well-curated and publicly available eQTL reference panels to create gene-expression prediction models, which may then be applied to \"impute\" genetically regulated gene expression (GREX) in large GWAS datasets. By testing for association between phenotype and GREX, rather than genotype, we hope to identify more biologically interpretable associations, and thus elucidate more of the genetic architecture of bipolar disorder.\n\nWe applied GREX prediction models for 13 brain regions (derived from CommonMind Consortium and GTEx eQTL reference panels) to 21,488 bipolar cases and 54,303 matched controls, constituting the largest transcriptomic imputation study of bipolar disorder (BPD) to date. Additionally, we analyzed three specific BPD subtypes, including 14,938 individuals with subtype 1 (BD-I), 3,543 individuals with subtype 2 (BD-II), and 1,500 individuals with schizoaffective subtype (SAB).\n\nWe identified 125 gene-tissue associations with BPD, of which 53 represent independent associations after FINEMAP analysis. 29/53 associations were novel; i.e., did not lie within 1Mb of a locus identified in the recent PGC-BD GWAS. We identified 37 independent BD-I gene-tissue associations (10 novel), 2 BD-II associations, and 2 SAB associations. Our BPD, BD-I and BD-II associations were significantly more likely to be differentially expressed in post-mortem brain tissue of BPD, BD-I and BD-II cases than we might expect by chance. Together with our pathway analysis, our results support long-standing hypotheses about bipolar disorder risk, including a role for oxidative stress and mitochondrial dysfunction, the post-synaptic density, and an enrichment of circadian rhythm and clock genes within our results.

genetics