bioRxiv Science⌕ Search

Biology subjects

Shahsavari, A.

Publications and source records attributed to Shahsavari, A..

3 recordsLinked to original sources

ClustAssess: tools for assessing the robustness of single-cell clustering

The transition from bulk to single-cell analyses refocused the computational challenges for high-throughput sequencing data-processing. The core of single-cell pipelines is partitioning cells and assigning cell-identities; extensive consequences derive from this step; generating robust and reproducible outputs is essential. From benchmarking established single-cell pipelines, we observed that clustering results critically depend on algorithmic choices (e.g. method, parameters) and technical details (e.g. random seeds). We present ClustAssess, a suite of tools for quantifying clustering robustness both within and across methods. The tools provide fine-grained information enabling (a) the detection of optimal number of clusters, (b) identification of regions of similarity (and divergence) across methods, (c) a data driven assessment of optimal parameter ranges. The aim is to assist practitioners in evaluating the robustness of cell-identity inference based on the partitioning, and provide information for choosing robust clustering methods and parameters. We illustrate its use on three case studies: a single-cell dataset of in-vivo hematopoietic stem and progenitors (10x Genomics scRNA-seq), in-vitro endoderm differentiation (SMART-seq), and multimodal in-vivo peripheral blood (10x RNA+ATAC). The additional checks offer novel viewpoints on clustering stability, and provide a framework for consistent decision-making on preprocessing, method choice, and parameters for clustering.

bioinformatics↗

The sum of two halves may be different from the whole. Effects of splitting sequencing samples across lanes.

The advances in high throughput sequencing (HTS) enabled the characterisation of biological processes at an unprecedented level of detail; the majority of hypotheses in molecular biology rely on analyses of HTS data. However, achieving increased robustness and reproducibility of results remains one of the main challenges. Although variability in results may be introduced at various stages, e.g. alignment, summarisation or detection of differences in expression, one source of variability was systematically omitted: the sequencing design which propagates through analyses and may introduce an additional layer of technical variation. We illustrate qualitative and quantitative differences arising from splitting samples across lanes, on bulk and single-cell sequencing. For bulk mRNAseq data, we focus on differential expression and enrichment analyses; for bulk ChIPseq data, we investigate the effect on peak calling, and peaks properties. At single-cell level, we concentrate on identifying cell subpopulations. We rely on markers used for assigning cell identities; both smartSeq and 10x data are presented. The observed reduction in the number of unique sequenced fragments reduces the level of detail on which the different prediction approaches depend. Further, the sequencing stochasticity adds in a weighting bias corroborated with variable sequencing depths and (yet unexplained) sequencing bias.

bioinformatics↗

Mapping the biogenesis of forward programmed megakaryocytes from induced pluripotent stem cells

Platelet deficiency, known as thrombocytopenia, can cause haemorrhage and is treated with platelet transfusions. We developed a system for the production of platelet precursor cells, megakaryocytes, from pluripotent stem cells. These cultures can be maintained for >100 days, implying culture renewal by megakaryocyte progenitors (MKPs). However, it is unclear whether the MKP state in vitro mirrors the state in vivo, and MKPs cannot be purified using conventional surface markers. We performed single cell RNA sequencing throughout in vitro differentiation and mapped each state to its equivalent in vivo. This enabled the identification of 5 surface markers which reproducibly purify MKPs, allowing us an insight into their transcriptional and epigenetic profiles. Finally, we performed culture optimisation, increasing MKP production. Altogether, this study has mapped parallels between the MKP states in vivo and in vitro and allowed the purification of MKPs, accelerating the progress of in vitro-derived transfusion products towards the clinic.

developmental biology↗