bioRxiv Science⌕ Search

Biology subjects

SO, H.-C.

Publications and source records attributed to SO, H.-C..

4 recordsLinked to original sources

SR2: Sparse Representation Learning for Scalable Single-cell RNA Sequencing Data Analysis

Single-cell RNA-sequencing (scRNA-seq) technology has been widely used to measure the transcriptome of cells in complex and heterogeneous systems. Integrative analysis of multiple scRNA-seq data can transform our understanding of various aspects of biology at the single-cell level. Many computational methods are proposed for data integration. However, few methods for scRNA-seq data integration explicitly model variation from heterogeneous biological conditions for interpretation. Modeling the variation helps understand the effect of biological conditions on complex biological systems. Our study proposes SR2 to capture gene expression patterns from heterogeneous biological conditions and discover cell identity simultaneously. Therefore, it can uncover the effect of biological conditions on the gene expression of cells and simultaneously achieve state-of-the-performance in cell identity discovery in our comprehensive comparison. Notably, SR2 is extended to model the effects of biological conditions on gene expression for cell populations, thus uncovering the effect of biological conditions on gene expression for cell populations and identifying putative condition-associated cell populations. To improve its scalability, we incorporate a batch-fitting strategy to ensure it is scalable to scRNA-seq data with arbitrary sample sizes. Moreover, the broad applicability of SR2 in biomedical studies has been demonstrated via applications. The complete package of SR2 is available at https://github.com/kai0511/SR2.

bioinformatics↗

INSIDER: Interpretable Sparse Matrix Decomposition for Bulk RNA Expression Data Analysis

RNA-Seq is widely used to capture transcriptome dynamics across tissues from different biological entities even across biological conditions, with the aim of understanding the contribution of gene activities to phenotypes of biosamples. However, due to variation from tissues and biological entities (or other biological conditions), joint analysis of bulk RNA expression profiles across multiple tissues from a number of biological entities to achieve the aim is hindered. Moreover, it is crucial to consider interactions between biological variables. For example, different brain disorders may affect brain regions heterogeneously. Thus, modeling the disorder-region interaction can shed light on the heterogeneity. To address these key challenges, we propose a general and flexible statistical framework based on matrix factorization, named INSIDER (https://github.com/kai0511/insider). INSIDER decomposes variation from different biological variables into a shared low-rank latent space. In particular, it considers interactions between biological variables and introduces the elastic net penalty to induce sparsity, thus facilitating interpretation. In the framework, the biological variables and interaction terms can be defined based on the research questions and study design. Besides, it enables us to compute the adjusted expression profiles for biological variables that control variation from other biological variables. Lastly, it allows various downstream analyses, such as clustering donors with donor representations, revealing development trajectory in its application to the BrainSpan data, and uncovering mechanisms underlying variables like phenotype and interactions between biological variables (e.g., phenotypes and tissues).

bioinformatics↗

Prediction of drug targets for specific diseases leveraging gene perturbation data: A machine learning approach

Identification of the correct targets is a key element for successful drug development. However, there are limited approaches for predicting drug targets for specific diseases using omics data, and few have leveraged expression profiles from gene perturbations. We present a novel computational target discovery approach based on machine learning(ML) models. ML models are first trained on drug-induced expression profiles, with outcomes defined as whether the drug treats the studied disease. The goal is to "learn" expression patterns associated with treatment. The fitted ML models were then applied to expression profiles from gene perturbations(over-expression[OE]/knockdown[KD]). We prioritized targets based on predicted probabilities from the ML model, which reflects treatment potential. The methodology was applied to predict targets for hypertension, diabetes mellitus(DM), rheumatoid arthritis(RA) and schizophrenia(SCZ). We validated our approach by evaluating whether the identified targets may re-discover known drug targets from an external database(OpenTargets). We indeed found evidence of significant enrichment across all diseases under study. Further literature search revealed that many candidates were supported by previous studies. For example, we predicted PSMB8 inhibition to be associated with treatment of RA, which was supported by a study showing PSMB8 inhibitors(PR-957) ameliorated experimental RA in mice. In conclusion, we propose a new ML approach to integrate expression profiles from drugs and gene perturbations and validated the framework. Our approach is flexible and may provide an independent source of information when prioritizing targets.

bioinformatics↗

Uncovering bi-directional causal relationships between plasma proteins and psychiatric disorders: A proteome-wide study leveraging GWAS summary data

Psychiatric disorders represent a major public health burden yet their etiologies remain poorly understood, and treatment advances are limited. In addition, there are no reliable biomarkers for diagnosis or progress monitoring. Here we performed a proteome-wide causal association study covering 3522 plasma proteins and 24 psychiatric traits or disorders, based on large-scale GWAS data and the principle of Mendelian randomization (MR). We have conducted ~95,000 MR analyses in total; to our knowledge, this is the most comprehensive study on the causal relationship between plasma proteins and psychiatric traits. The analysis was bi-directional: we studied how proteins may affect psychiatric disorder risks, but also looked into how psychiatric traits/disorders may be causal risk factors for changes in protein levels. We also performed a variety of additional analysis to prioritize protein-disease associations, including HEIDI test for distinguishing functional association from linkage, analysis restricted to cis- acting variants and replications in independent datasets from the UK Biobank. Based on the MR results, we constructed directed networks linking proteins, drugs and different psychiatric traits, hence shedding light on their complex relationships and drug repositioning opportunities. Interestingly, many top proteins were related to inflammation or immune functioning. The full results were also made available online in searchable databases. In conclusion, identifying proteins causal to disease development have important implications on drug discovery or repurposing. Findings from this study may also guide the development of blood-based biomarkers for the prediction or diagnosis of psychiatric disorders, as well as assessment of disease progression or recovery.

genomics↗