bioRxiv ScienceSearch

Biology subjects

Su, S.

Publications and source records attributed to Su, S..

8 recordsLinked to original sources

scRNA-seq mixology: towards better benchmarking of single cell RNA-seq protocols and analysis methods

Single cell RNA sequencing (scRNA-seq) technology has undergone rapid development in recent years, bringing with it new challenges in data processing and analysis. This has led to an explosion of tailored analysis methods for scRNA-seq to address various biological questions. However, the current lack of gold-standard benchmarking datasets makes it difficult for researchers to evaluate the performance of the many methods. Here, we designed and carried out a realistic benchmark experiment that included mixtures of single cells or pseudo-cells created by sampling admixtures of cells or RNA from 3 distinct cancer cell lines. Altogether we generated 10 datasets using a combination of droplet and plate-based scRNA-seq protocols, with varying data quality, population heterogeneity and noise levels. Using these benchmark datasets, we compared different protocols, evaluated the spike-in standard and multiple data analysis methods for tasks ranging from normalization and imputation, to clustering, trajectory analysis and data integration. Evaluation of methods across multiple datasets revealed some that performed well in general and others that suited specific situations. Our dataset and analysis provide a comprehensive comparison framework for benchmarking most popular scRNA-seq analysis tasks.

bioinformatics

SIS-seq, a molecular ‘time machine’, connects single cell fate with gene programs

Conventional single cell RNA-seq methods are destructive, such that a given cell cannot also then be tested for fate and function, without a time machine. Here, we develop a clonal method SIS-seq, whereby single cells are allowed to divide, and progeny cells are assayed separately in SISter conditions; some for fate, others by RNA-seq. By cross-correlating progenitor gene expression with mature cell fate within a clone, and doing this for many clones, we can identify the earliest gene expression signatures of dendritic cell subset development. SIS-seq could be used to study other populations harboring clonal heterogeneity, including stem, reprogrammed and cancer cells to reveal the transcriptional origins of fate decisions.

systems biology

A data-driven approach to characterising intron signal in RNA-seq data

RNA-seq datasets can contain millions of intron reads per sequenced library that are typically removed from downstream analysis. Only reads overlapping annotated exons are considered to be informative since mature mRNA is assumed to be the major component sequenced, especially when examining poly(A) RNA samples. In this paper, we demonstrate that intron reads are informative and that pre-mRNA is the major source of intron signal. Making use of pre-mRNA signal, our index method combines differential expression analyses from intron and exon counts to categorise changes observed in each count set, giving additional genes with evidence of transcriptional changes when compared to a classic approach. Considering the importance of intron retention in some biological systems, another novel method, superintronic, looks for evidence of intron retention after accounting for the presence of pre-mRNA signal. The results presented here overcomes deficiencies and biases in previous works related to intron reads by exploring multiple sources for intron reads simultaneously using a data-driven approach, and provides a broad overview into how intron reads can be utilised in relation to multiple aspects of transcriptional biology.

bioinformatics

ALV-J and REV synergistically activate a new oncogene of KIAA1199 via NF-κB and EGFR signaling regulated by miR-147

The tumorigenesis is the result of the accumulation of multiple oncogenes and tumor suppressor genes changes. Co-infection of avian leucosis virus subgroup J (ALV-J) and reticuloendotheliosis virus (REV), as two oncogenic retroviruses, showed synergistic pathogenic effects characterized by enhanced tumor initiation and progression. The molecular mechanism underlying synergistic effects of ALV-J and REV on the neoplasia remains unclear. Here, we found co-infection of ALV-J and REV enhanced the ability of virus infection, increased viral life cycle, maintained cell survival and enhanced tumor formation. We combined the high-throughput proteomic readout with a large-scale miRNA screening to identify which molecules are involved in the synergism. Our results revealed co-infection of ALV-J and REV activated a latent oncogene of KIAA1199 and inhibited the expression of tumor suppressor miR-147. Further, enhanced KIAA1199, down-regulated miR-147, activated NF-{kappa}B and EGFR were demonstrated in co-infected tissues and tumor. Mechanistically, we showed ALV-J and REV synergistically enhanced KIAA1199 by activation of NF-{kappa}B and EGFR signalling pathway, and the suppression of tumor suppressor miR-147 was contributed to maintain the NF-{kappa}B/KIAA1199/EGFR pathway crosstalk by targeting the 3UTR region sequences of NF-{kappa}B p50 and KIAA1199. Our results contributed to the understanding of the molecular mechanisms of viral synergistic tumorgenesis, which provided the evidence that suggested the synergistic actions of two retroviruses could result in activation of latent pro-oncogenes.\n\nAuthor summaryThe tumorigenesis is the result of the accumulation of multiple oncogenes and tumor suppressor genes changes. Co-infection with ALV-J and REV showed synergistic pathogenic effects characterized by enhanced tumor progression, however, the molecular mechanism on the neoplasia remains unclear. Our results revealed co-infection of ALV-J and REV promotes tumorigenesis by both induction of a latent oncogene of KIAA1199 and suppression of the expression of tumor suppressor miR-147. Mechanistic studies revealed that ALV-J and REV synergistically enhance KIAA1199 by activation of NF-{kappa}B and EGFR signalling pathway, and the suppression of tumor suppressor miR-147 was contributed to maintain the NF-{kappa}B/KIAA1199/EGFR pathway crosstalk by targeting the 3UTR region sequences of NF-{kappa}B p50 and KIAA1199. These results provided the evidence that suggested the synergistic actions of two retroviruses could result in activation of latent pro-oncogenes, indicating the potential preventive target and predictive factor for ALV-J and REV induced tumorigenesis.

molecular biology

scPipe: a flexible data preprocessing pipeline for single-cell RNA-sequencing data

Single-cell RNA sequencing (scRNA-seq) technology allows researchers to profile the transcriptomes of thousands of cells simultaneously. Protocols that incorpo-rate both designed and random barcodes have greatly increased the throughput of scRNA-seq, but give rise to a more complex data structure. There is a need for new tools that can handle the various barcoding strategies used by different protocols and exploit this information for quality assessment at the sample-level and provide effective visualization of these results in preparation for higher-level analyses.\n\nTo this end, we developed scPipe, a R/Bioconductor package that integrates barcode demultiplexing, read alignment, UMI-aware gene-level quantification and quality control of raw sequencing data generated by multiple 3-prime-end sequencing protocols that include CEL-seq, MARS-seq, Chromium 10X and Drop-seq. scPipe produces a count matrix that is essential for downstream analysis along with an HTML report that summarises data quality. These results can be used as input for downstream analyses including normalization, visualization and statistical testing. scPipe performs this processing in a few simple R commands, promoting reproducible analysis of single-cell data that is compatible with the emerging suite of scRNA-seq analysis tools available in R/Bioconductor. The scPipe R package is available for download from https://www.bioconductor.org/packages/scPipe.

bioinformatics

Structural Basis For The Specific Recognition Of DSR By The YTH Domain Containing Protein Mmi1

Meiosis is one of the most dramatic differentiation programs accompanied by a striking change in gene expression profiles, whereas a number of meiosis-specific transcripts are expressed untimely in mitotic cells. The entry of meiosis will be blocked as the accumulation of meiosis-specific mRNAs during the mitotic cell in fission yeast Schizosaccharomyces pombe. A YTH domain containing protein Mmi1 was identified as a pivotal effector in a post-transcriptional event termed selective elimination of meiosis-specific mRNAs, Mmi1 can recognize and bind a class of meiosis-specific transcripts expressed inappropriately in mitotic cells, which contain a conservative motif called DSR as a mark to remove them in cooperation with nuclear exosomes. Here we report the 1.6 [A] resolution crystal structure of the YTH domain of Mmi1 binds to high-affinity RNA targets r(A1U2U3A4A5A6C7A8) containing DSR core motif. Our structure observations, supported by site-directed mutations of key residues illustrate the mechanism for specific recognition of DSR-RNA by Mmi1. Moreover, different from other YTH domain family proteins, Mmi1 YTH domain has a distinctive function although it has a similar fold as other ones.

biochemistry

Molecular Basis For The Specific And Multivariate Recognitions Of RNA Substrates By Human hnRNPA2/B1

Human hnRNPA2/B1 is an RNA-binding protein that plays important roles in a variety of biological processes, from mRNA maturation, trafficking and translation to regulation of gene expression mediated by long non-coding RNAs and microRNAs. hnRNPA2/B1 contains two RNA recognition motifs (RRM) that provide sequence-specific recognition of widespread RNA substrates including recently reported m6A-containing motifs. Here we determined the first crystal structures of tandem RRM domains of hnRNPA2/B1 in complex with various RNA substrates. Our structures reveal that hnRNPA2/B1 can bind two RNA elements in an antiparallel fashion with a sequence preference for AGG and UAG by RRM1 and RRM2, respectively, suggesting an RNA matchmaker mechanism during the hnRNPA2/B1 function. However, our combined studies did not observe specific binding of m6A by either the RRM domains or the full-length hnRNPA2/B1, implying that the \"reader\" function of hnRNPA2/B1 may adopt an unknown mechanism that remains to be characterized.

biochemistry

Glimma: interactive graphics for gene expression analysis

MotivationSummary graphics for RNA-sequencing and microarray gene expression analyses may contain upwards of tens of thousands of points. Details about certain genes or samples of interest are easily obscured in such dense summary displays. Incorporating interactivity into summary plots would enable additional information to be displayed on demand and facilitate intuitive data exploration.\n\nResultsThe open-source Glimma package creates interactive graphics for exploring gene expression analysis with a few simple R commands. It extends popular plots found in the limma package, such as multi-dimensional scaling plots and mean-difference plots, to allow individual data points to be queried and additional annotation information to be displayed upon hovering or selecting particular points. It also offers links between plots so that more information can be revealed on demand. Glimma is widely applicable, supporting data analyses from a number of well established Bioconductor workflows (limma, edgeR and DESeq2) and uses D3/JavaScript to produce HTML pages with interactive displays that enable more effective data exploration by end-users. Results from Glimma can be easily shared between bioinformaticians and biologists, enhancing reporting capabilities while maintaining reproducibility.\n\nAvailability and ImplementationThe Glimma R package is available from http://bioconductor.org/packages/devel/bioc/html/Glimma.html.

bioinformatics