bioRxiv ScienceSearch

Biology subjects

Marioni, J. C.

Publications and source records attributed to Marioni, J. C..

11 recordsLinked to original sources

Staged developmental mapping and X chromosome transcriptional dynamics during mouse spermatogenesis

Understanding male fertility requires an in-depth characterisation of spermatogenesis, the developmental process by which male gametes are generated. Spermatogenesis occurs continuously throughout a males reproductive window and involves a complex sequence of developmental steps, both of which make this process difficult to decipher at the molecular level. To overcome this, we transcriptionally profiled single cells from multiple distinct stages during the first wave of spermatogenesis, where the most mature germ cell type is known. This naturally enriches for spermatogonia and somatic cell types present at very low frequencies in adult testes. Our atlas, available as a shiny app (https://marionilab.cruk.cam.ac.uk/SpermatoShiny), allowed us to reconstruct the three main processes of spermatogenesis: spermatogonial differentiation, meiosis, and spermiogenesis. Additionally, we profiled the chromatin changes associated with meiotic silencing of the X chromosome, revealing a set of genes specifically and strongly repressed by H3K9me3 in the spermatocyte stage, but which escape post-meiotic silencing in spermatids.

developmental biology

Robust expression variability testing reveals heterogeneous T cell responses

Cell-to-cell transcriptional variability in otherwise homogeneous cell populations plays a crucial role in tissue function and development. Single-cell RNA sequencing can characterise this variability in a transcriptome-wide manner. However, technical variation and the confounding between variability and mean expression estimates hinders meaningful comparison of expression variability between cell populations. To address this problem, we introduce a novel analysis approach that extends the BASiCS statistical framework to derive a residual measure of variability that is not confounded by mean expression. Moreover, we introduce a new and robust procedure for quantifying technical noise in experiments where technical spike-in molecules are not available. We illustrate how our method provides biological insight into the dynamics of cell-to-cell expression variability, highlighting a synchronisation of the translational machinery in immune cells upon activation. Additionally, our approach identifies new patterns of variability across CD4+ T cell differentiation.

bioinformatics

Multi-Omics factor analysis disentangles heterogeneity in blood cancer

Multi-omic studies promise the improved characterization of biological processes across molecular layers. However, methods for the unsupervised integration of the resulting heterogeneous datasets are lacking. We present Multi-Omics Factor Analysis (MOFA), a computational method for discovering the principal sources of variation in multi-omic datasets. MOFA infers a set of (hidden) factors that capture biological and technical sources of variability. It disentangles axes of heterogeneity that are shared across multiple modalities and those specific to individual data modalities. The learnt factors enable a variety of downstream analyses, including identification of sample subgroups, data imputation, and the detection of outlier samples. We applied MOFA to a cohort of 200 patient samples of chronic lymphocytic leukaemia, profiled for somatic mutations, RNA expression, DNA methylation and ex-vivo drug responses. MOFA identified major dimensions of disease heterogeneity, including immunoglobulin heavy chain variable region status, trisomy of chromosome 12 and previously underappreciated drivers, such as response to oxidative stress. In a second application, we used MOFA to analyse single-cell multiomics data, identifying coordinated transcriptional and epigenetic changes along cell differentiation.

bioinformatics

Placozoans are eumetazoans related to Cnidaria

The phylogenetic placement of the morphologically simple placozoans is crucial to understanding the evolution of complex animal traits. Here, we examine the influence of adding new genomes from placozoans to a large dataset designed to study the deepest splits in the animal phylogeny. Using site-heterogeneous substitution models, we show that it is possible to obtain strong support, in both amino acid and reduced-alphabet matrices, for either a sister-group relationship between Cnidaria and Placozoa, or for Cnidaria and Bilateria (=Planulozoa), also seen in most published work to date, depending on the orthologues selected to construct the matrix. We demonstrate that a majority of genes show evidence of compositional heterogeneity, and that the support for Planulozoa can be assigned to this source of systematic error. In interpreting this placozoan-cnidarian clade, we caution against a peremptory reading of placozoans as secondarily reduced forms of little relevance to broader discussions of early animal evolution.

zoology

Detection and removal of barcode swapping in single-cell RNA-seq data

Barcode swapping results in the mislabeling of sequencing reads between multiplexed samples on the new patterned flow cell Illumina sequencing machines. This may compromise the validity of numerous genomic assays, especially for single-cell studies where many samples are routinely multiplexed together. The severity and consequences of barcode swapping for single-cell transcriptomic studies remain poorly understood. We have used two statistical approaches to robustly quantify the fraction of swapped reads in each of two plate-based single-cell RNA sequencing datasets. We found that approximately 2.5% of reads were mislabeled between samples on the HiSeq 4000 machine, which is lower than previous reports. We observed no correlation between the swapped fraction of reads and the concentration of free barcode across plates. Furthermore, we have demonstrated that barcode swapping may generate complex but artefactual cell libraries in droplet-based single-cell RNA sequencing studies. To eliminate these artefacts, we have developed an algorithm to exclude individual molecules that have swapped between samples in 10X Genomics experiments, exploiting the combinatorial complexity present in the data. This permits the continued use of cutting-edge sequencing machines for droplet-based experiments while avoiding the confounding effects of barcode swapping.

genomics

Whole-body single-cell sequencing of the Platynereis larva reveals a subdivision into apical versus non-apical tissues

Animal bodies comprise a diverse array of tissues and cells. To characterise cellular identities across an entire body, we have compared the transcriptomes of single cells randomly picked from dissociated whole larvae of the marine annelid Platynereis dumerilii1-4. We identify five transcriptionally distinct groups of differentiated cells that are spatially coherent, as revealed by spatial mapping5. Besides somatic musculature, ciliary bands and midgut, we find a group of cells located at the apical tip of the animal, comprising sensory-peptidergic neurons, and another group composed of non-apical neural and epidermal cells covering the rest of the body. These data establish a basic subdivision of the larval body surface into molecularly defined apical versus non-apical tissues, and support the evolutionary conservation of the apical nervous system as a distinct part of the bilaterian brain6.

evolutionary biology

Correcting batch effects in single-cell RNA sequencing data by matching mutual nearest neighbours.

The presence of batch effects is a well-known problem in experimental data analysis, and single- cell RNA sequencing (scRNA-seq) is no exception. Large-scale scRNA-seq projects that generate data from different laboratories and at different times are rife with batch effects that can fatally compromise integration and interpretation of the data. In such cases, computational batch correction is critical for eliminating uninteresting technical factors and obtaining valid biological conclusions. However, existing methods assume that the composition of cell populations are either known or the same across batches. Here, we present a new strategy for batch correction based on the detection of mutual nearest neighbours in the high-dimensional expression space. Our approach does not rely on pre-defined or equal population compositions across batches, only requiring that a subset of the population be shared between batches. We demonstrate the superiority of our approach over existing methods on a range of simulated and real scRNA-seq data sets. We also show how our method can be applied to integrate scRNA-seq data from two separate studies of early embryonic development.

bioinformatics

Joint Profiling Of Chromatin Accessibility, DNA Methylation And Transcription In Single Cells

Parallel single-cell sequencing protocols represent powerful methods for investigating regulatory relationships, including epigenome-transcriptome interactions. Here, we report a novel single-cell method for parallel chromatin accessibility, DNA methylation and transcriptome profiling. scNMT-seq (single-cell nucleosome, methylation and transcription sequencing) uses a GpC methyltransferase to label open chromatin followed by bisulfite and RNA sequencing. We validate scNMT-seq by applying it to differentiating mouse embryonic stem cells, finding links between all three molecular layers and revealing dynamic coupling between epigenomic layers during differentiation.

genomics

Assessing The Reliability Of Spike-In Normalization For Analyses Of Single-Cell RNA Sequencing Data

By profiling the transcriptomes of individual cells, single-cell RNA sequencing provides unparalleled resolution to study cellular heterogeneity. However, this comes at the cost of high technical noise, including cell-specific biases in capture efficiency and library generation. One strategy for removing these biases is to add a constant amount of spike-in RNA to each cell, and to scale the observed expression values so that the coverage of spike-in RNA is constant across cells. This approach has previously been criticized as its accuracy depends on the precise addition of spike-in RNA to each sample, and on similarities in behaviour (e.g., capture efficiency) between the spike-in and endogenous transcripts. Here, we perform mixture experiments using two different sets of spike-in RNA to quantify the variance in the amount of spike-in RNA added to each well in a plate-based protocol. We also obtain an upper bound on the variance due to differences in behaviour between the two spike-in sets. We demonstrate that both factors are small contributors to the total technical variance and have only minor effects on downstream analyses such as detection of highly variable genes and clustering. Our results suggest that spike-in normalization is reliable enough for routine use in single-cell RNA sequencing data analyses.

bioinformatics

Flipping between Polycomb repressed and active transcriptional states introduces noise in gene expression

Polycomb repressive complexes (PRCs) are important histone modifiers, which silence gene expression, yet there exists a subset of PRC-bound genes actively transcribed by RNA polymerase II (RNAPII). It is likely that the role of PRC is to dampen expression of these PRC-active genes. However, it is unclear how this flipping between chromatin states alters the kinetics of transcriptional burst size and frequency relative to genes with exclusively activating marks. To investigate this, we integrate histone modifications and RNAPII states derived from bulk ChIP-seq data with single-cell RNA-sequencing data. We find that PRC-active genes have a greater cell-to-cell variation in expression than active genes with the same mean expression levels, and validate these results by knockout experiments. We also show that PRC-active genes are clustered on chromosomes in both two and three dimensions, and interactions with active enhancers promote a stabilization of gene expression noise. These findings provide new insights into how chromatin regulation modulates stochastic gene expression and transcriptional bursting, with implications for regulation of pluripotency and development.

genomics

Scalable latent-factor models applied to single-cell RNA-seq data separate biological drivers from confounding effects

Single-cell RNA-sequencing (scRNA-seq) allows heterogeneity in gene expression levels to be studied in large populations of cells. Such heterogeneity can arise from both technical and biological factors, thus making decomposing sources of variation extremely difficult. We here describe a computationally efficient model that uses prior pathway annotation to guide inference of the biological drivers underpinning the heterogeneity. Moreover, we jointly update and improve gene set annotation and infer factors explaining variability that fall outside the existing annotation. We validate our method using simulations, which demonstrate both its accuracy and its ability to scale to large datasets with up to 100,000 cells. Moreover, through applications to real data we show that our model can robustly decompose scRNA-seq datasets into interpretable components and facilitate the identification of novel sub-populations.

bioinformatics