bioRxiv ScienceSearch

Biology subjects

Mahfouz, A.

Publications and source records attributed to Mahfouz, A..

4 recordsLinked to original sources

Conserved cell types with divergent features between human and mouse cortex

Elucidating the cellular architecture of the human neocortex is central to understanding our cognitive abilities and susceptibility to disease. Here we applied single nucleus RNA-sequencing to perform a comprehensive analysis of cell types in the middle temporal gyrus of human cerebral cortex. We identify a highly diverse set of excitatory and inhibitory neuronal types that are mostly sparse, with excitatory types being less layer-restricted than expected. Comparison to a similar mouse cortex single cell RNA-sequencing dataset revealed a surprisingly well-conserved cellular architecture that enables matching of homologous types and predictions of human cell type properties. Despite this general conservation, we also find extensive differences between homologous human and mouse cell types, including dramatic alterations in proportions, laminar distributions, gene expression, and morphology. These species-specific features emphasize the importance of directly studying human brain.

neuroscience

Single-cell isoform RNA sequencing (ScISOr-Seq) across thousands of cells reveals isoforms of cerebellar cell types.

Full-length isoform sequencing has advanced our knowledge of isoform biology1-11. However, apart from applying full-length isoform sequencing to very few single cells12,13, isoform sequencing has been limited to bulk tissue, cell lines, or sorted cells. Single splicing events have been described for <=200 single cells with great statistical success14,15, but these methods do not describe full-length mRNAs. Single cell short-read 3 sequencing has allowed identification of many cell sub-types16-23, but full-length isoforms for these cell types have not been profiled. Using our new method of single-cell-isoform-RNA-sequencing (ScISOr-Seq) we determine isoform-expression in thousands of individual cells from a heterogeneous bulk tissue (cerebellum), without specific antibody-fluorescence activated cell sorting. We elucidate isoform usage in high-level cell types such as neurons, astrocytes and microglia and finer sub-types, such as Purkinje cells and Granule cells, including the combination patterns of distant splice sites6-9,24,25, which for individual molecules requires long reads. We produce an enhanced genome annotation revealing cell-type specific expression of known and 16,872 novel (with respect to mouse Gencode version 10) isoforms (see isoformatlas.com).\n\nScISOr-Seq describes isoforms from >1,000 single cells from bulk tissue without cell sorting by leveraging two technologies in three steps: In step one, we employ microfluidics to produce amplified full-length cDNAs barcoded for their cell of origin. This cDNA is split into two pools: one pool for 3 sequencing to measure gene expression (step 2) and another pool for long-read sequencing and isoform expression (step 3). In step two, short-read 3-sequencing provides molecular counts for each gene and cell, which allows clustering cells and assigning a cell type using cell-type specific markers. In step three, an aliquot of the same cDNAs (each barcoded for the individual cell of origin) is sequenced using Pacific Biosciences (\"PacBio\")1,2,4,5,26 or Oxford Nanopore3. Since these long reads carry the single-cell barcodes identified in step two, one can determine the individual cell from which each long read originates. Since most single cells are assigned to a named cluster, we can also assign the cells cluster name (e.g. \"Purkinje cell\" or \"astrocyte\") to the long read in question (Fig 1A) - without losing the cell of origin of each long read.\n\nO_FIG O_LINKSMALLFIG WIDTH=180 HEIGHT=200 SRC=\"FIGDIR/small/364950_fig1.gif\" ALT=\"Figure 1\">\nView larger version (66K):\norg.highwire.dtl.DTLVardef@4df0cdorg.highwire.dtl.DTLVardef@fc4beborg.highwire.dtl.DTLVardef@1dc485forg.highwire.dtl.DTLVardef@1138a3e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1:C_FLOATNO (A) Outline of our ScISOr-Seq approach. (B) TSNE-plot depicting cell clusters, marker genes and names given to clusters, including: Bergman glia (BG), External granule cell layer neurons (EGL), Internal granule cell layer and other neurons in the interior of the cerebellum (IGL), two clusters of Purkinje cell layer neurons (PCL), oligodendrocyte progenitor cells (OPCs), Atoh1+ neuronal progenitors, Ptf1a+ neuronal progenitors and other neuronal progenitors (NPCs) (C) In-situ hybridization images from the Allen Brain Atlas depicting expression of marker genes in specific layers. (D) Expression patterns of selected marker genes across cell types.\n\nC_FIG

molecular biology

Predicting cell types in single cell mass cytometry data

MotivationMass cytometry (CyTOF) is a valuable technology for high-dimensional analysis at the single cell level. Identification of different cell populations is an important task during the data analysis. Many clustering tools can perform this task, however, they are time consuming, often involve a manual step, and lack reproducibility when new data is included in the analysis. Learning cell types from an annotated set of cells solves these problems. However, currently available mass cytometry classifiers are either complex, dependent on prior knowledge of the cell type markers during the learning process, or can only identify canonical cell types.\n\nResultsWe propose to use a Linear Discriminant Analysis (LDA) classifier to automatically identify cell populations in CyTOF data. LDA shows comparable results with two state-of-the-art algorithms on four benchmark datasets and also outperforms a non-linear classifier such as the k-nearest neighbour classifier. To illustrate its scalability to large datasets with deeply annotated cell subtypes, we apply LDA to a dataset of ~3.5 million cells representing 57 cell types. LDA has high performance on abundant cell types as well as the majority of rare cell types, and provides accurate estimates of cell type frequencies. Further incorporating a rejection option, based on the estimated posterior probabilities, allows LDA to identify cell types that were not encountered during training. Altogether, reproducible prediction of cell type compositions using LDA opens up possibilities to analyse large cohort studies based on mass cytometry data.\n\nAvailabilityImplementation is available on GitHub (https://github.com/tabdelaal/CyTOF-Linear-Classifier).\n\nContacta.mahfouz@lumc.nl

bioinformatics

A structural equation model for imaging genetics using spatial transcriptomics

Alzheimers disease is a neurodegenerative disorder that causes changes in the structure of the brain, observable with MRI scans, and that has a strong heritable component, reflected in the DNA. Imaging genetics deals with such relationships between genetic variation and imaging variables, often in a disease context. The complex relationships between brain volumes and genetic variants have been explored both with dimension reduction methods and model based approaches. However, these models usually do not make use of the extensive knowledge of the spatio-anatomical patterns of gene activity. We present a method for integrating genetic markers (single nucleotide polymorphisms) and imaging features, which is based on a causal model and, at the same time, uses the power of dimension reduction. We use structural equation models to find latent variables that explain brain volume changes in a disease context, and which are in turn affected by genetic variants. We make use of publicly available spatial transcriptome data from the Allen Human Brain Atlas to specify the model structure, which reduces noise and improves interpretability. The model is tested in a simulation setting, and applied on a case study of the Alzheimers Disease Neuroimaging Initiative.

bioinformatics