bioRxiv Science⌕ Search

Biology subjects

Grabski, I. N.

Publications and source records attributed to Grabski, I. N..

2 recordsLinked to original sources

Differentially methylated regions and methylation QTLs for teen depression and early puberty in the Fragile Families Child Wellbeing Study

AO_SCPLOWBSTRACTC_SCPLOWThe Fragile Families Child Wellbeing Study (FFCWS) is a longitudinal cohort of ethnically diverse and primarily low socioeconomic status children and their families in the U.S. Here, we analyze DNA methylation data collected from 748 FFCWS participants in two waves of this study, corresponding to participant ages 9 and 15. Our primary goal is to leverage the DNA methylation data from these two time points to study methylation associated with two key traits in adolescent health that are over-represented in these data: Early puberty and teen depression. We first identify differentially methylated regions (DMRs) for depression and early puberty. We then identify DMRs for the interaction effects between these two conditions and age by including interaction terms in our regression models to understand how age-related changes in methylation are influenced by depression or early puberty. Next, we identify methylation quantitative trait loci (meQTLs) using genotype data from the participants. We also identify meQTLs with epistatic effects with depression and early puberty. We find enrichment of our interaction meQTLs with functional categories of the genome that contribute to the heritability of co-morbid complex diseases. We replicate our meQTLs in data from the GoDMC study. This work leverages the important focus of the FFCWS data on disadvantaged children to shed light on the methylation states associated with teen depression and early puberty, and on how genetic regulation of methylation is affected in adolescents with these two conditions.

genomics↗

Probabilistic gene expression signatures identify cell-types from single cell RNA-seq data

AO_SCPLOWBSTRACTC_SCPLOWSingle-cell RNA sequencing (scRNA-seq) quantifies gene expression for individual cells in a sample, which allows distinct cell-type populations to be identified and characterized. An important step in many scRNA-seq analysis pipelines is the annotation of cells into known cell-types. While this can be achieved using experimental techniques, such as fluorescence-activated cell sorting, these approaches are impractical for large numbers of cells. This motivates the development of data-driven cell-type annotation methods. We find limitations with current approaches due to the reliance on known marker genes or from overfitting because of systematic differences between studies or batch effects. Here, we present a statistical approach that leverages public datasets to combine information across thousands of genes, uses a latent variable model to define cell-type-specific barcodes and account for batch effect variation, and probabilistically annotates cell-type identity. The barcoding approach also provides a new way to discover marker genes. Using a range of datasets, including those generated to represent imperfect real-world reference data, we demonstrate that our approach substantially outperforms current reference-based methods, in particular when predicting across studies. Our approach also demonstrates that current approaches based on unsupervised clustering lead to false discoveries related to novel cell-types.

genomics↗