bioRxiv ScienceSearch

Biology subjects

Gut, M.

Publications and source records attributed to Gut, M..

10 recordsLinked to original sources

De novo assembly and annotation of the larval transcriptome of two spadefoot toads widely divergent in developmental rate

Introduction Introduction Methods Results and Discussion Conclusion Data and materials References Most amphibian species exhibit a complex life-cycle including two or more life stages separated by an ontogenetic switch point such as hatching or metamorphosis. Adaptations to divergent environments can require the modification of the timing of such switch points and the relative investment in growth and differentiation between subsequent stages [1]. Such alterations of developmental trajectories, however, often have substantial repercussions at several organismal levels, from physiology to morphology and even genomic structure. Adaptive divergence in developmental rate tracking aquatic habitats of different duration in spadefoot toads is a well-known example of this. Spadefoot toads from Europe and ...

genomics

Single cell expression analysis uncouples transdifferentiation and reprogramming

Many somatic cell types are plastic, having the capacity to convert into other specialized cells (transdifferentiation)(1) or into induced pluripotent stem cells (iPSCs, reprogramming)(2) in response to transcription factor over-expression. To explore what makes a cell plastic and whether these different cell conversion processes are coupled, we exposed bone marrow derived pre-B cells to two different transcription factor overexpression protocols that efficiently convert them either into macrophages or iPSCs and monitored the two processes over time using single cell gene expression analysis. We found that even in these highly efficient cell fate conversion systems, cells differ in both their speed and path of transdifferentiation and reprogramming. This heterogeneity originatesin two starting pre-B cell subpopulations,large pre-BII and the small pre-BII cells they normally differentiate into. The large cells transdifferentiate slowly but exhibit a high efficiency of iPSC reprogramming. In contrast, the small cells transdifferentiate rapidly but are highly resistant to reprogramming. Moreover, the large B cells induce a stronger transient granulocyte/macrophage progenitor (GMP)-like state, while the small B cells undergo a more direct conversion to the macrophage fate. The large cells are cycling and exhibit high Myc activity whereas the small cells are Myc low and mostly quiescent. The observed heterogeneity of the two cell conversion processes can therefore be traced to two closely related cell types in the starting population that exhibit different types of plasticity. These data show that a somatic cells propensity for either transdifferentiation and reprogramming can be uncoupled.\n\nOne sentence summarySingle cell transcriptomics of cell conversions

developmental biology

Selective single molecule sequencing and assembly of a human Y chromosome of African origin

Mammalian Y chromosomes are often neglected from genomic analysis. Due to their inherent assembly difficulties, high repeat content, and large ampliconic regions1, only a handful of species have their Y chromosome properly characterized. To date, just a single human reference quality Y chromosome, of European ancestry, is available due to a lack of accessible methodology2-5. To facilitate the assembly of such complicated genomic territory, we developed a novel strategy to sequence native, unamplified flow sorted DNA on a MinION nanopore sequencing device. Our approach yields a highly continuous and complete assembly of the first human Y chromosome of African origin. It constitutes a significant improvement over comparable previous methods, increasing continuity by more than 800%6, thus allowing a chromosome scale analysis of human Y chromosomes. Sequencing native DNA also allows to take advantage of the nanopore signal data to detect epigenetic modifications in situ7. This approach is in theory generalizable to any species simplifying the assembly of extremely large and repetitive genomes.

genomics

pheno-seq - linking 3D phenotypes of clonal tumor spheroids to gene expression

3D-culture systems have advanced cancer modeling by reflecting physiological characteristics of in-vivo tissues, but our understanding of functional intratumor heterogeneity including visual phenotypes and underlying gene expression is still limited. Single-cell RNA-sequencing is the method of choice to dissect transcriptional tumor cell heterogeneity in an unbiased way, but this approach is limited in correlating gene expression with contextual cellular phenotypes.\n\nTo link morphological features and gene expression in 3D-culture systems, we present pheno-seq for integrated high-throughput imaging and transcriptomic profiling of clonal tumor spheroids. Specifically, we identify characteristic EMT expression signatures that are associated with invasive growth behavior in a 3D breast cancer model. Additionally, pheno-seq determined transcriptional programs containing lineage-specific markers that can be linked to heterogeneous proliferative capacity in a patient-derived 3D model of colorectal cancer. Finally, we provide evidence that pheno-seq identifies morphology-specific genes that are missed by scRNA-seq and inferred single-cell regulatory states without acquiring additional single cell expression profiles. We anticipate that directly linking molecular features with patho-phenotypes of cancer cells will improve the understanding of intratumor heterogeneity and consequently be useful for translational research.

genomics

Partially methylated domains are hypervariable in breast cancer and fuel widespread CpG island hypermethylation

Global loss of DNA methylation and CpG island (CGI) hypermethylation are regarded as key epigenomic aberrations in cancer. Global loss manifests itself in partially methylated domains (PMDs) which can extend up to megabases. However, the distribution of PMDs within and between tumor types, and their effects on key functional genomic elements including CGIs are poorly defined. Using whole genome bisulfite sequencing (WGBS) of breast cancers, we comprehensively show that loss of methylation in PMDs occurs in a large fraction of the genome and represents the prime source of variation in DNA methylation. PMDs are hypervariable in methylation level, size and distribution, and display elevated mutation rates. They impose intermediate DNA methylation levels incognizant of functional genomic elements including CGIs, underpinning a CGI methylator phenotype (CIMP). However, significant repression effects on cancer-genes are negligible as tumor suppressor genes are generally excluded from PMDs. The genomic distribution of PMDs reports tissue-of-origin of different cancers and may represent tissue-specific silent regions of the genome, which tolerate instability at the epigenetic, transcriptomic and genetic level.

cancer biology

Epigenomic and functional dynamics of human bone marrow myeloid differentiation to mature blood neutrophils

Neutrophils are short-lived blood cells that play a critical role in host defense against infections. To better comprehend neutrophil functions and their regulation, we provide a complete epigenetic and functional overview of their differentiation stages from bone marrow-residing progenitors to mature circulating cells. Integration of epigenetic and transcriptome dynamics reveals an enforced regulation of differentiation, through cellular functions such as: release of proteases, respiratory burst, cell cycle regulation and apoptosis. We observe an early establishment of the cytotoxic capability, whilst the signaling components that activate antimicrobial mechanisms are transcribed at later stages, outside the bone marrow, thus preventing toxic effects in the bone marrow niche. Altogether, these data reveal how the developmental dynamics of the epigenetic landscape orchestrate the daily production of large number of neutrophils required for innate host defense and provide a comprehensive overview of the epigenomes of differentiating human neutrophils.\n\nKey pointsO_LIDynamic acetylation enforces human neutrophil progenitor differentiation.\nC_LI\n\nO_LINeutrophils cytotoxic capability is established early at the (pro)myelocyte stage.\nC_LI\n\nO_LICoordinated signaling component expression prevents unwanted toxic effects to the bone marrow niche.\nC_LI

immunology

DNA methylation oscillation defines classes of enhancers

Understanding the regulatory landscape of human cells requires the integration of genomic and epigenomic maps, capturing combinatorial levels of cell type-specific and invariant activity states.\n\nHere, we segmented whole-genome bisulfite sequencing-derived methylomes into consecutive blocks of co-methylation (COMETs) to obtain spatial variation patterns of DNA methylation (DNAm oscillations) integrated with histone modifications and promoter-enhancer interactions derived from promoter capture Hi-C (PCHi-C) sequencing of the same purified blood cells.\n\nMapping DNAm oscillations onto regulatory genome annotation revealed that enhancers are enriched for DNAm hyper-oscillations (>30-fold), where multiple machine learning models support DNAm as predictive of enhancer location. Based on this analysis, we report overall predictive power of 99% for DNAm oscillations, 77.3% for DNaseI, 41% for CGIs, 20% for UMRs and 0% for LMRs, demonstrating the power of DNAm oscillations over other methods for enhancer prediction. Methylomes of activated and non-activated CD4+ T cells indicate that DNAm oscillations exist in both states irrespective of activation; hence they can be used to determine the location of latent enhancers.\n\nOur approach advances the identification of tissue-specific regulatory elements and outperforms previous approaches defining enhancer classes based on DNA methylation.

genomics

Comparative analysis of neutrophil and monocyte epigenomes

Neutrophils and monocytes provide a first line of defense against infections as part of the innate immune system. Here we report the integrated analysis of transcriptomic and epigenetic landscapes for circulating monocytes and neutrophils with the aim to enable downstream interpretation and functional validation of key regulatory elements in health and disease. We collected RNA-seq data, ChIP-seq of six histone modifications and of DNA methylation by bisulfite sequencing at base pair resolution from up to 6 individuals per cell type. Chromatin segmentation analyses suggested that monocytes have a higher number of cell-specific enhancer regions (4-fold) compared to neutrophils. This highly plastic epigenome is likely indicative of the greater differentiation potential of monocytes into macrophages, dendritic cells and osteoclasts. In contrast, most of the neutrophil-specific features tend to be characterized by repressed chromatin, reflective of their status as terminally differentiated cells. Enhancers were the regions where most of differences in DNA methylation between cells were observed, with monocyte-specific enhancers being generally hypomethylated. Monocytes show a substantially higher gene expression levels than neutrophils, in line with epigenomic analysis revealing that gene more active elements in monocytes. Our analyses suggest that the overexpression of c-Myc in monocytes and its binding to monocyte-specific enhancers could be an important contributor to these differences. Altogether, our study provides a comprehensive epigenetic chart of chromatin states in primary human neutrophils and monocytes, thus providing a valuable resource for studying the regulation of the human innate immune system.

genomics

bigSCale: An Analytical Framework for Big-Scale Single-Cell Data

Single-cell RNA sequencing significantly deepened our insights into complex tissues and latest techniques are capable processing ten-thousands of cells simultaneously. With bigSCale, we provide an analytical framework being scalable to analyze millions of cells, addressing challenges of future large datasets. Unlike previous methods, bigSCale does not constrain data to fit an a priori-defined distribution and instead uses an accurate numerical model of noise. We evaluated the performance of bigSCale using a biological model of aberrant gene expression in patient derived neuronal progenitor cells and simulated datasets, which underlined its speed and accuracy in differential expression analysis. We further applied bigSCale to analyze 1.3 million cells from the mouse developing forebrain. Herein, we identified rare populations, such as Reelin positive Cajal-Retzius neurons, for which we determined a previously not recognized heterogeneity associated to distinct differentiation stages, spatial organization and cellular function. Together, bigSCale presents a perfect solution to address future challenges of large single-cell datasets.\n\nExtended AbstractSingle-cell RNA sequencing (scRNAseq) significantly deepened our insights into complex tissues by providing high-resolution phenotypes for individual cells. Recent microfluidic-based methods are scalable to ten-thousands of cells, enabling an unbiased sampling and comprehensive characterization without prior knowledge. Increasing cell numbers, however, generates extremely big datasets, which extends processing time and challenges computing resources. Current scRNAseq analysis tools are not designed to analyze datasets larger than from thousands of cells and often lack sensitivity and specificity to identify marker genes for cell populations or experimental conditions. With bigSCale, we provide an analytical framework for the sensitive detection of population markers and differentially expressed genes, being scalable to analyze millions of single cells. Unlike other methods that use simple or mixture probabilistic models with negative binomial, gamma or Poisson distributions to handle the noise and sparsity of scRNAseq data, bigSCale does not constrain the data to fit an a priori-defined distribution. Instead, bigSCale uses large sample sizes to estimate a highly accurate and comprehensive numerical model of noise and gene expression. The framework further includes modules for differential expression (DE) analysis, cell clustering and population marker identification. Moreover, a directed convolution strategy allows processing of extremely large data sets, while preserving the transcript information from individual cells.\n\nWe evaluate the performance of bigSCale using a biological model for reduced or elevated gene expression levels. Specifically, we perform scRNAseq of 1,920 patient derived neuronal progenitor cells from Williams-Beuren and 7q11.23 microduplication syndrome patients, harboring a deletion or duplication of 7q11.23, respectively. The affected region contains 28 genes whose transcriptional levels vary in line with their allele frequency. BigSCale detects expression changes with respect to cells from a healthy donor and outperforms other methods for single-cell DE analysis in sensitivity. Simulated data sets, underline the performance of bigSCale in DE analysis as it is faster and more sensitive and specific than other methods. The probabilistic model of cell-distances within bigSCale is further suitable for unsupervised clustering and the identification of cell types and subpopulations. Using bigSCale, we identify all major cell types of the somatosensory cortex and hippocampus analyzing 3,005 cells from adult mouse brains. Remarkably, we increase the number of cell population specific marker genes 4-6-fold compared to the original analysis and, moreover, define markers of higher order cell types. These include CD90 (Thy1), a neuronal surface receptor, potentially suitable for isolating intact neurons from complex brain samples.\n\nTo test its applicability for large data sets, we apply bigSCale on scRNAseq data from 1.3 million cells derived from the pallium of the mouse developing forebrain (E18, 10x Genomics). Our directed down-sampling strategy accumulates transcript counts from cells with similar transcriptional profiles into index cell transcriptomes, thereby defining cellular clusters with improved resolution. Accordingly, index cell clusters provide a rich resource of marker genes for the main brain cell types and less frequent subpopulations. Our analysis of rare populations includes poorly characterized developmental cell types, such as neuron progenitors from the subventricular zone and neocortical Reelin positive neurons known as Cajal-Retzius (CR) cells. The latter represent a transient population which regulates the laminar formation of the developing neocortex and whose malfunctioning causes major neurodevelopmental disorders like autism or schizophrenia. Most importantly, index cell cluster can be deconvoluted to individual cell level for targeted analysis of populations of interest. Through decomposition of Reelin positive neurons, we determined a previously not recognized heterogeneity among CR cells, which we could associate to distinct differentiation stages as well as spatial and functional differences in the developing mouse brain. Specifically, subtypes of CR cells identified by bigSCale express different compositions of NMDA, AMPA and glycine receptor subunits, pointing to subpopulations with distinct membrane properties. Furthermore, we found Cxcl12, a chemokine secreted by the meninges and regulating the tangential migration of CR cells, to be also expressed in CR cells located in the marginal zone of the neocortex, indicating a self-regulated migration capacity.\n\nTogether, bigSCale presents a perfect solution for the processing and analysis of scRNAseq data from millions of single cells. Its speed and sensitivity makes it suitable to the address future challenges of large single-cell data sets.

genomics

Framework For Quality Assessment Of Whole Genome, Cancer Sequences

Working with cancer whole genomes sequenced over a period of many years in different sequencing centres requires a validated framework to compare the quality of these sequences. The Pan-Cancer Analysis of Whole Genomes (PCAWG) of the International Cancer Genome Consortium (ICGC), a project a cohort of over 2800 donors provided us with the challenge of assessing the quality of the genome sequences. A non-redundant set of five quality control (QC) measurements were assembled and used to establish a star rating system. These QC measures reflect known differences in sequencing protocol and provide a guide to downstream analyses of these whole genome sequences. The resulting QC measures also allowed for exclusion samples of poor quality, providing researchers within PCAWG, and when the data is released for other researchers, a good idea of the sequencing quality. For a researcher wishing to apply the QC measures for their data we provide a Docker Container of the software used to calculate them. We believe that this is an effective framework of quality measures for whole genome, cancer sequences, which will be a useful addition to analytical pipelines, as it has to the PCAWG project.

genomics