bioRxiv Science⌕ Search

Biology subjects

Sonthalia, S.

Publications and source records attributed to Sonthalia, S..

3 recordsLinked to original sources

CellCover Defines Conserved Cell Types and Temporal Progression in scRNA-seq Data across Mammalian Neocortical Development

1Definition of cell classes across the tissues of living organisms is central in the analysis of growing atlases of single-cell RNA sequencing (scRNA-seq) data across biomedicine. Marker genes for cell classes are most often defined by differential expression (DE) methods that serially assess individual genes across landscapes of diverse cells. This serial approach has been extremely useful, but is limited because it ignores possible redundancy or complementarity across genes that can only be captured by analyzing multiple genes simultaneously. Interrogating binarized expression data, we aim to identify discriminating panels of genes that are specific to, not only enriched in, individual cell types. To efficiently explore the vast space of possible marker panels, leverage the large number of cells often sequenced, and overcome zero-inflation in scRNA-seq data, we propose viewing marker gene panel selection as a variation of the "minimal set-covering problem" in combinatorial optimization. Using scRNA-seq data from blood and brain tissue, we show that this new method, CellCover, performs as good or better than DE and other methods in defining cell-type discriminating gene panels, while reducing gene redundancy and capturing cell-class-specific signals that are distinct from those defined by DE methods. Transfer learning experiments across mouse, primate, and human data demonstrate that CellCover identifies markers of conserved cell classes in neocortical neurogenesis, as well as developmental progression in both progenitors and neurons. Exploring markers of human outer radial glia (oRG, or basal RG) across mammals, we show that transcriptomic elements of this key cell type in the expansion of the human cortex likely appeared in gliogenic precursors of the rodent before the full program emerged in neurogenic cells of the primate lineage. We have assembled the public datasets we use in this report within the NeMO Analytics multi-omic data exploration environment [1], where the expression of individual genes (NeMO: Individual genes in cortex and NeMO: Individual genes in blood) and marker gene panels (NeMO: Telley 3 CellCover Panels, NeMO: Telley 12 CellCover Panels, NeMO: Sorted Brain Cell CellCover Panels, and NeMO: Blood 34 CellCover Panels) can be freely explored without coding expertise. CellCover is available in CellCover R and CellCover Python. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=67 SRC="FIGDIR/small/535943v6_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@893301org.highwire.dtl.DTLVardef@173a6baorg.highwire.dtl.DTLVardef@1c70e7eorg.highwire.dtl.DTLVardef@1888ab3_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

Structured Joint Decomposition (SJD) identifies conserved molecular dynamics across collections of biologically related multi-omics data matrices

It is necessary to develop exploratory tools to learn from the unprecedented volume of high-dimensional multi-omic data currently being produced across the field of biomedicine. We have developed an R package, Structured Joint Decomposition (SJD), which identifies components of variation that are shared across multiple matrices. The approach focuses specifically on variation across the samples/cells within each dataset while incorporating biologist-defined hierarchical structure among input experiments that can span in vivo and in vitro systems, multi-omic data modalities, and species. SJD enables the definition of molecular variation that is conserved across systems, those that are shared within subsets of studies, and elements unique to individual matrices. We have included functions to simplify the construction and visualization of highly complex in silico experiments involving many diverse multi-omic matrices from multiple species. Here we apply SJD to decompose four RNA-seq experiments focused on neurogenesis in the neocortex. The public datasets used in this analysis are at NeMO Analytics and can be explored at the individual gene level or using the conserved transcriptomic dynamics in mammalian neurogenesis that we define here. The SJD R package and tutorial can be found at https://chuansite.github.io/SJD. Contact: hzchenhuan@gmail.com; ccolant1@jhmi.edu [carlocolantuoni.org]

bioinformatics↗

Insights for disease modeling from single cell transcriptomics of iPSC-derived neurons and astrocytes across differentiation time and co-culture

Trans-differentiation of human induced pluripotent stem cells into neurons via Ngn2-induction (hiPSC-N) has become an efficient system to quickly generate neurons for disease modeling and in vitro assay development, a significant advance from previously used neoplastic and other cell lines. Recent single-cell interrogation of Ngn2-induced neurons however, has revealed some similarities to unexpected neuronal lineages. Similarly, a straightforward method to generate hiPSC derived astrocytes (hiPSC-A) for the study of neuropsychiatric disorders has also been described. Here we examine the homogeneity and similarity of hiPSC-N and hiPSC-A to their in vivo counterparts, the impact of different lengths of time post Ngn2 induction on hiPSC-N (15 or 21 days) and of hiPSC-N / hiPSC-A co-culture. Leveraging the wealth of existing public single-cell RNA-seq (scRNA-seq) data in Ngn2-induced neurons and in vivo data from the developing brain, we provide perspectives on the lineage origins and maturation of hiPSC-N and hiPSC-A. While induction protocols in different labs produce consistent cell type profiles, both hiPSC-N and hiPSC-A show significant heterogeneity and similarity to multiple in vivo cell fates, and both more precisely approximate their in vivo counterparts when co-cultured. Gene expression data from the hiPSC-N show enrichment of genes linked to schizophrenia (SZ) and autism spectrum disorders (ASD) as has been previously shown for neural stem cells and neurons. These overrepresentations of disease genes are strongest in our system at early times (day 15) in Ngn2-induction/maturation of neurons, when we also observe the greatest similarity to early in vivo excitatory neurons. We have assembled this new scRNA-seq data along with the public data explored here as an integrated biologist-friendly web-resource for researchers seeking to understand this system more deeply: nemoanalytics.org/p?l=DasEtAlNGN2&g=PRPH.

genetics↗