bioRxiv ScienceSearch

Biology subjects

Nguyen, Q.

Publications and source records attributed to Nguyen, Q..

8 recordsLinked to original sources

Generation of human neural retina transcriptome atlas by single cell RNA sequencing

The retina is a highly specialized neural tissue that senses light and initiates image processing. Although the functional organisation of specific cells within the retina has been well-studied, the molecular profile of many cell types remains unclear in humans. To comprehensively profile cell types in the human retina, we performed single cell RNA-sequencing on 20,009 cells obtained post-mortem from three donors and compiled a reference transcriptome atlas. Using unsupervised clustering analysis, we identified 18 transcriptionally distinct clusters representing all known retinal cells: rod photoreceptors, cone photoreceptors, Muller glia cells, bipolar cells, amacrine cells, retinal ganglion cells, horizontal cells, retinal astrocytes and microglia. Notably, our data captured molecular profiles for healthy and early degenerating rod photoreceptors, and revealed a novel role of MALAT1 in putative rod degeneration. We also demonstrated the use of this retina transcriptome atlas to benchmark pluripotent stem cell-derived cone photoreceptors and an adult Muller glia cell line. This work provides an important reference with unprecedented insights into the transcriptional landscape of human retinal cells, which is fundamental to our understanding of retinal biology and disease.

systems biology

scPred: Single cell prediction using singular value decomposition and machine learning classification

Single-cell RNA sequencing has enabled the characterization of highly specific cell types in many human tissues, as well as both primary and stem cell-derived cell lines. An important facet of these studies is the ability to identify the transcriptional signatures that define a cell type or state. In theory, this information can be used to classify an unknown cell based on its transcriptional profile; and clearly, the ability to accurately predict a cell type and any pathologic-related state will play a critical role in the early diagnosis of disease and decisions around the personalized treatment for patients. Here we present a new generalizable method (scPred) for prediction of cell type(s), using a combination of unbiased feature selection from a reduced-dimension space, and machine-learning classification. scPred solves several problems associated with the identification of individual gene feature selection, and is able to capture subtle effects of many genes, increasing the overall variance explained by the model, and correspondingly improving the prediction accuracy. We validate the performance of scPred by performing experiments to classify tumor versus non-tumor epithelial cells in gastric cancer, then using independent molecular techniques (cyclic immunohistochemistry) to confirm our prediction, achieving an accuracy of classifying the disease state of individual cells of 99%. Moreover, we apply scPred to scRNA-seq data from pancreatic tissue, colorectal tumor biopsies, and circulating dendritic cells, and show that scPred is able to classify cell subtypes with an accuracy of 96.1-99.2%. Collectively, our results demonstrate the utility of scPred as a single cell prediction method that can be used for a wide variety of applications. The generalized method is implemented in software available here: https://github.com/IMB-Computational-Genomics-Lab/scPred/

genomics

Detection of HPV E7 transcription at single-cell resolution in epidermis

Persistent human papillomavirus (HPV) infection is responsible for at least 5% of human malignancies. Most HPV-associated cancers are initiated by the HPV16 genotype, as confirmed by detection of integrated HPV DNA in cells of oral and anogenital epithelial cancers. However, single-cell RNA-sequencing (scRNA-seq) may enable prediction of HPV involvement in carcinogenesis at other sites. We conducted scRNA-seq on keratinocytes from a mouse transgenic for the E7 gene of HPV16, and showed sensitive and specific detection of HPV16-E7 mRNA, predominantly in basal keratinocytes. We showed that increased E7 mRNA copy number per cell was associated with increased expression of E7 induced genes. This technique enhances detection of viral transcripts in solid tissue and may clarify possible linkage of HPV infection to development of squamous cell carcinoma.

cell biology

Cardiac directed differentiation using small molecule Wnt modulation at single-cell resolution

Differentiation into diverse cell lineages requires the orchestration of gene regulatory networks guiding diverse cell fate choices. Utilizing human pluripotent stem cells, we measured expression dynamics of 17,718 genes from 43,168 cells across five time points over a thirty day time-course of in vitro cardiac-directed differentiation. Unsupervised clustering and lineage prediction algorithms were used to map fate choices and transcriptional networks underlying cardiac differentiation. We leveraged this resource to identify strategies for controlling in vitro differentiation as it occurs in vivo. HOPX, a non-DNA binding homeodomain protein essential for heart development in vivo was identified as dys-regulated in in vitro derived cardiomyocytes. Utilizing genetic gain and loss of function approaches, we dissect the transcriptional complexity of the HOPX locus and identify the requirement of hypertrophic signaling for HOPX transcription in hPSC-derived cardiomyocytes. This work provides a single cell dissection of the transcriptional landscape of cardiac differentiation for broad applications of stem cells in cardiovascular biology.

developmental biology

Determining cell fate specification and genetic contribution to cardiac disease risk in hiPSC-derived cardiomyocytes at single cell resolution

The majority of genetic loci underlying common disease risk act through changing genome regulation, and are routinely linked to expression quantitative trait loci, where gene expression is measured using bulk populations of mature cells. A crucial step that is missing is evidence of variation in the expression of these genes as cells progress from a pluripotent to mature state. This is especially important for cardiovascular disease, as the majority of cardiac cells have limited properties for renewal postneonatal. To investigate the dynamic changes in gene expression across the cardiac lineage, we generated RNA-sequencing data captured from 43,168 single cells progressing through in vitro cardiac-directed differentiation from pluripotency. We developed a novel and generalized unsupervised cell clustering approach and a machine learning method for prediction of cell transition. Using these methods, we were able to reconstruct the cell fate choices as cells transition from a pluripotent state to mature cardiomyocytes, uncovering intermediate cell populations that do not progress to maturity, and distinct cell trajectories that terminate in cardiomyocytes that differ in their contractile forces. Second, we identify new gene markers that denote lineage specification and demonstrate a substantial increase in their utility for cell identification over current pluripotent and cardiogenic markers. By integrating results from analysis of the single cell lineage RNA-sequence data with population-based GWAS of cardiovascular disease and cardiac tissue eQTLs, we show that the pathogenicity of disease-associated genes is highly dynamic as cells transition across their developmental lineage, and exhibit variation between cell fate trajectories. Through the integration of single cell RNA-sequence data with population-scale genetic data we have identified genes significantly altered at cell specification events providing insights into a context-dependent role in cardiovascular disease risk. This study provides a valuable data resource focused on in vitro cardiomyocyte differentiation to understand cardiac disease coupled with new analytical methods with broad applications to single-cell data.

genomics

ascend: R package for analysis of single cell RNA-seq data

Summaryascend is an R package comprised of fast, streamlined analysis functions optimized to address the statistical challenges of single cell RNA-seq. The package incorporates novel and established methods to provide a flexible framework to perform filtering, quality control, normalization, dimension reduction, clustering, differential expression and a wide-range of plotting. ascend is designed to work with scRNA-seq data generated by any high-throughput platform, and includes functions to convert data objects between software packages.\n\nAvailabilityThe R package and associated vignettes are freely available at https://github.com/IMB-Computational-Genomics-Lab/ascend.\n\nContactjoseph.powell@uq.edu.au\n\nSupplementary informationAn example dataset is available at ArrayExpress, accession number E-MTAB-6108

bioinformatics

Single Cell RNA Sequencing of stem cell-derived retinal ganglion cells.

We used human embryonic stem cell-derived retinal ganglion cells (RGCs) to characterize the transcriptome of 1,174 cells at the single cell level. The human embryonic stem cell line BRN3B-mCherry A81-H7 was differentiated to RGCs using a guided differentiation approach. Cells were harvested at day 36 and subsequently prepared for single cell RNA sequencing. Our data indicates the presence of three distinct subpopulations of cells, with various degrees of maturity. One cluster of 288 cells upregulated genes involved in axon guidance together with semaphorin interactions, cell-extracellular matrix interactions and ECM proteoglycans, suggestive of a more mature phenotype.

genomics

Single-Cell Transcriptome Sequencing Of 18,787 Human Induced Pluripotent Stem Cells Identifies Differentially Primed Subpopulations

Heterogeneity of cell states represented in pluripotent cultures have not been described at the transcriptional level. Since gene expression is highly heterogeneous between cells, single-cell RNA sequencing can be used to identify how individual pluripotent cells function. Here, we present results from the analysis of single-cell RNA sequencing data from 18,787 individual WTC CRISPRi human induced pluripotent stem cells. We developed an unsupervised clustering method, and through this identified four subpopulations distinguishable on the basis of their pluripotent state including: a core pluripotent population (48.3%), proliferative (47.8%), early-primed for differentiation (2.8%) and late-primed for differentiation (1.1%). For each subpopulation we were able to identify the genes and pathways that define differences in pluripotent cell states. Our method identified four transcriptionally distinct predictor gene sets comprised of 165 unique genes that denote the specific pluripotency states; and using these sets, we developed a multigenic machine learning prediction method to accurately classify single cells into each of the subpopulations. Compared against a set of established pluripotency markers, our method increases prediction accuracy by 10%, specificity by 20%, and explains a substantially larger proportion of deviance (up to 3-fold) from the prediction model. Finally, we developed an innovative method to predict cells transitioning between subpopulations, and support our conclusions with results from two orthogonal pseudotime trajectory methods.

genomics