bioRxiv ScienceSearch

Biology subjects

Shannon, P.

Publications and source records attributed to Shannon, P..

3 recordsLinked to original sources

Reproducible big data science: A case study in continuous FAIRness

Big biomedical data create exciting opportunities for discovery, but make it difficult to capture analyses and outputs in forms that are findable, accessible, interoperable, and reusable (FAIR). In response, we describe tools that make it easy to capture, and assign identifiers to, data and code throughout the data lifecycle. We illustrate the use of these tools via a case study involving a multi-step analysis that creates an atlas of putative transcription factor binding sites from terabytes of ENCODE DNase I hypersensitive sites sequencing data. We show how the tools automate routine but complex tasks, capture analysis algorithms in understandable and reusable forms, and harness fast networks and powerful cloud computers to process data rapidly, all without sacrificing usability or reproducibility--thus ensuring that big data are not hard-to-(re)use data. We compare and contrast our approach with other approaches to big data analysis and reproducibility.

bioinformatics

Atlas of Transcription Factor Binding Sites from ENCODE DNase Hypersensitivity Data Across 27 Tissue Types

There is intense interest in mapping the tissue-specific binding sites of transcription factors in the human genome to reconstruct gene regulatory networks and predict functions for non-coding genetic variation. DNase-seq footprinting provides a means to predict genome-wide binding sites for hundreds of transcription factors (TFs) simultaneously. However, despite the public availability of DNase-seq data for hundreds of samples, there is neither a unified analytical workflow nor a publicly accessible database providing the locations of footprints across all available samples. Here, we implemented a workflow for uniform processing of footprints using two state-of-the-art footprinting algorithms: Wellington and HINT. Our workflow scans the footprints generated by these algorithms for 1,530 sequence motifs to predict binding sites for 1,515 human transcription factors. We applied our workflow to detect footprints in 192 DNase-seq experiments from ENCODE spanning 27 human tissues. This collection of footprints describes an expansive landscape of potential TF occupancy. At thresholds optimized through machine learning, we report high-quality footprints covering 9.8% of the human genome. These footprints were enriched for true positive TF binding sites as defined by ChIP-seq peaks, as well as for genetic variants associated with changes in gene expression. Integrating our footprint atlas with summary statistics from genome-wide association studies revealed that risk for neuropsychiatric traits was enriched specifically at highly-scoring footprints in human brain, while risk for immune traits was enriched specifically at highly-scoring footprints in human lymphoblasts. Our cloud-based workflow is available at github.com/globusgenomics/genomics-footprint and a database with all footprints and TF binding site predictions are publicly available at http://data.nemoarchive.org/other/grant/sament/sament/footprint_atlas.

bioinformatics

Genome-scale transcriptional regulatory network models of psychiatric and neurodegenerative disorders

Genetic and genomic studies suggest an important role for transcriptional regulatory changes in brain diseases, but roles for specific transcription factors (TFs) remain poorly understood. We integrated human brain-specific DNase I footprinting and TF-gene co-expression to reconstruct a transcriptional regulatory network (TRN) model for the human brain, predicting the brain-specific binding sites and target genes for 741 TFs. We used this model to predict core TFs involved in psychiatric and neurodegenerative diseases. Our results suggest that disease-related transcriptomic and genetic changes converge on small sets of disease-specific regulators, with distinct networks underlying neurodegenerative vs. psychiatric diseases. Core TFs were frequently implicated in a disease through multiple mechanisms, including differential expression of their target genes, disruption of their binding sites by disease-associated SNPs, and associations of the genetic loci encoding these TFs with disease risk. We validated our models predictions through systematic comparison to publicly available ChIP-seq and TF perturbation studies and through experimental studies in primary human neural stem cells. Combined genetic and transcriptional evidence supports roles for neuronal and microglia-enriched, MEF2C-regulated networks in Alzheimers disease; an oligodendrocyte-enriched, SREBF1-regulated network in schizophrenia; and a neural stem cell and astrocyte-enriched, POU3F2-regulated network in bipolar disorder. We provide our models of brain-specific TF binding sites and target genes as a resource for network analysis of brain diseases.

genetics