bioRxiv ScienceSearch

Biology subjects

Schadt, E. E.

Publications and source records attributed to Schadt, E. E..

5 recordsLinked to original sources

Epigenomic landscape of the human pathogen Clostridium difficile

Clostridioides difficile is a leading cause of health care-associated infections. Although significant progress has been made in the understanding of its genome, the epigenome of C. difficile and its functional impact has not been systematically explored. Here, we performed the first comprehensive DNA methylome analysis of C. difficile using 36 human isolates and observed great epigenomic diversity. We discovered an orphan DNA methyltransferase with a well-defined specificity whose corresponding gene is highly conserved across our dataset and in all ~300 global C. difficile genomes examined. Inactivation of the methyltransferase gene negatively impacted sporulation, a key step in C. difficile disease transmission, consistently supported by multi-omics data, genetic experiments, and a mouse colonization model. Further experimental and transcriptomic analysis also suggested that epigenetic regulation is associated with cell length, biofilm formation, and host colonization. These findings open up a new epigenetic dimension to characterize medically relevant biological processes in this critical pathogen. This work also provides a set of methods for comparative epigenomics and integrative analysis, which we expect to be broadly applicable to bacterial epigenomics studies.

genomics

Functional Interpretation of Genetic Variants Using Deep Learning Predicts Impact on Epigenome

Identifying causal variants underling disease risk and adoption of personalized medicine are currently limited by the challenge of interpreting the functional consequences of genetic variants. Predicting the functional effects of disease-associated protein-coding variants is increasingly routine. Yet the vast majority of risk variants are non-coding, and predicting the functional consequence and prioritizing variants for functional validation remains a major challenge. Here we develop a deep learning model to accurately predict locus-specific signals from four epigenetic assays using only DNA sequence as input. Given the predicted epigenetic signal from DNA sequence for the reference and alternative alleles at a given locus, we generate a score of the predicted epigenetic consequences for 438 million variants. These impact scores are assay-specific, are predictive of allele-specific transcription factor binding and are enriched for variants associated with gene expression and disease risk. Nucleotide-level functional consequence scores for non-coding variants can refine the mechanism of known causal variants, identify novel risk variants and prioritize downstream experiments.

genomics

Discovering genetic interactions bridging pathways in genome-wide association studies

Genetic interactions have been reported to underlie phenotypes in a variety of systems, but the extent to which they contribute to complex disease in humans remains unclear. In principle, genome-wide association studies (GWAS) provide a platform for detecting genetic interactions, but existing methods for identifying them from GWAS data tend to focus on testing individual locus pairs, which undermines statistical power. Importantly, the global genetic networks mapped for a model eukaryotic organism revealed that genetic interactions often connect genes between compensatory functional modules in a highly coherent manner. Taking advantage of this expected structure, we developed a computational approach called BridGE that identifies pathways connected by genetic interactions from GWAS data. Applying BridGE broadly, we discovered significant interactions in Parkinsons disease, schizophrenia, hypertension, prostate cancer, breast cancer, and type 2 diabetes. Our novel approach provides a general framework for mapping complex genetic networks underlying human disease from genome-wide genotype data.

genetics

Integrative analyses of splicing in the aging brain: role in susceptibility to Alzheimer’s Disease

We use deep sequencing to identify sources of variation in mRNA splicing in the dorsolateral prefrontal cortex (DLFPC) of 450 subjects from two prospective cohort studies of aging. Hundreds of aberrant pre-mRNA splicing events are reproducibly associated with Alzheimers Disease (AD). We also generate a catalog of splicing quantitative trait loci (sQTL) effects in the human cortex: splicing of 3,198 genes is influenced by genetic variation. sQTLs are enriched among those variants influencing DNA methylation and histone acetylation. In assessing known AD loci, we report that altered splicing is the mechanism for the effects of the PICALM, CLU, and PTK2B susceptibility alleles. Further, we leverage our sQTL catalog to identify genes whose aberrant splicing is associated with AD and mediated by genetics. This transcriptome-wide association study identified 21 genes with significant associations, many of which are found in AD GWAS loci, but 8 are in novel AD loci, including FUS, which is a known amyotrophic lateral sclerosis (ALS) gene. This highlights an intriguing shared genetic architecture that is further elaborated by the convergence of old and new AD genes in autophagy-lysosomal-related pathways already implicated in AD and other neurodegenerative diseases. Overall, this study of the aging brains transcriptome provides evidence that dysregulation of mRNA splicing is a feature of AD and is, in some genetically-driven cases, causal.

genomics

A Novel Nasal Brush-based Classifier of Asthma Identified by Machine Learning Analysis of Nasal RNA Sequence Data

Asthma is a common, under-diagnosed disease affecting all ages. We sought to identify a nasal brush-based classifier of mild/moderate asthma. 190 subjects with mild/moderate asthma and controls underwent nasal brushing and RNA sequencing of nasal samples. A machine learning-based pipeline identified an asthma classifier consisting of 90 genes interpreted via an L2-regularized logistic regression classification model. This classifier performed with strong predictive value and sensitivity across eight test sets, including (1) a test set of independent asthmatic and control subjects profiled by RNA sequencing (positive and negative predictive values of 1.00 and 0.96, respectively; AUC of 0.994), (2) two independent case-control cohorts of asthma profiled by microarray, and (3) five cohorts with other respiratory conditions (allergic rhinitis, upper respiratory infection, cystic fibrosis, smoking), where the classifier had a low to zero misclassification rate. Following validation in large, prospective cohorts, this classifier could be developed into a nasal biomarker of asthma.

systems biology