bioRxiv Science⌕ Search

Biology subjects

ZHAO, K.

Publications and source records attributed to ZHAO, K..

2 recordsLinked to original sources

SR2: Sparse Representation Learning for Scalable Single-cell RNA Sequencing Data Analysis

Single-cell RNA-sequencing (scRNA-seq) technology has been widely used to measure the transcriptome of cells in complex and heterogeneous systems. Integrative analysis of multiple scRNA-seq data can transform our understanding of various aspects of biology at the single-cell level. Many computational methods are proposed for data integration. However, few methods for scRNA-seq data integration explicitly model variation from heterogeneous biological conditions for interpretation. Modeling the variation helps understand the effect of biological conditions on complex biological systems. Our study proposes SR2 to capture gene expression patterns from heterogeneous biological conditions and discover cell identity simultaneously. Therefore, it can uncover the effect of biological conditions on the gene expression of cells and simultaneously achieve state-of-the-performance in cell identity discovery in our comprehensive comparison. Notably, SR2 is extended to model the effects of biological conditions on gene expression for cell populations, thus uncovering the effect of biological conditions on gene expression for cell populations and identifying putative condition-associated cell populations. To improve its scalability, we incorporate a batch-fitting strategy to ensure it is scalable to scRNA-seq data with arbitrary sample sizes. Moreover, the broad applicability of SR2 in biomedical studies has been demonstrated via applications. The complete package of SR2 is available at https://github.com/kai0511/SR2.

bioinformatics↗

INSIDER: Interpretable Sparse Matrix Decomposition for Bulk RNA Expression Data Analysis

RNA-Seq is widely used to capture transcriptome dynamics across tissues from different biological entities even across biological conditions, with the aim of understanding the contribution of gene activities to phenotypes of biosamples. However, due to variation from tissues and biological entities (or other biological conditions), joint analysis of bulk RNA expression profiles across multiple tissues from a number of biological entities to achieve the aim is hindered. Moreover, it is crucial to consider interactions between biological variables. For example, different brain disorders may affect brain regions heterogeneously. Thus, modeling the disorder-region interaction can shed light on the heterogeneity. To address these key challenges, we propose a general and flexible statistical framework based on matrix factorization, named INSIDER (https://github.com/kai0511/insider). INSIDER decomposes variation from different biological variables into a shared low-rank latent space. In particular, it considers interactions between biological variables and introduces the elastic net penalty to induce sparsity, thus facilitating interpretation. In the framework, the biological variables and interaction terms can be defined based on the research questions and study design. Besides, it enables us to compute the adjusted expression profiles for biological variables that control variation from other biological variables. Lastly, it allows various downstream analyses, such as clustering donors with donor representations, revealing development trajectory in its application to the BrainSpan data, and uncovering mechanisms underlying variables like phenotype and interactions between biological variables (e.g., phenotypes and tissues).

bioinformatics↗