bioRxiv ScienceSearch

Biology subjects

Sobhan, M.

Publications and source records attributed to Sobhan, M..

3 recordsLinked to original sources

Quantifying Intratumor Heterogeneity by Key Genes Selected Using Concrete Autoencoder

The tumor cell population in cancer tissue has distinct molecular characteristics and exhibits different phenotypes, thus, resulting in different subpopulations. This phenomenon is known as Intratumor Heterogeneity (ITH), a major contributor to drug resistance, poor prognosis, etc. Therefore, quantifying the levels of ITH in cancer patients is essential, and many algorithms do so in different ways, using different types of omics data. DEPTH (Deviating gene Expression Profiling Tumor Heterogeneity) is the latest algorithm that uses transcriptomic data to evaluate the ITH score. It shows promising performance, has strong similarity with six other algorithms and has an advantage over two algorithms that uses the same type of data (tITH, sITH). However, it has a major drawback since it uses expression values of all the genes ([~]20K genes) in quantifying ITH levels. We hypothesize that a subset of key genes is sufficient to quantify the ITH level. To prove our hypothesis, we developed a deep learning-based computational framework using unsupervised Concrete Autoencoder (CAE) to select a set of cancer-specific key genes that can be used to evaluate the ITH score. For the experiment, we used gene expression profile data of tumor cohorts of breast, kidney, and lung cancer from the TCGA repository. Using multi-run CAE, we selected three sets of key genes, each set related to breast, kidney, and lung tumor cohorts. For the three cancers stated and three molecular subtypes of lung cancer, we calculated the ITH level using all genes and key genes selected by CAE and performed a side-by-side comparison. We could reach similar conclusions for survival and prognostic outcomes based on ITH scores derived from all genes and the sets of key genes. Additionally, for subtypes of lung cancer, the comparative distribution of ITH scores derived from all and key genes remains similar. Based on these observations, it can be stated that a subset of key genes, instead of all genes, is sufficient for ITH quantification. Our results also showed that many key genes are prognostically significant, which can be used as possible therapeutic targets.

bioinformatics

Molecular mimicry between Spike and human thrombopoietin may induce thrombocytopenia in COVID-19

SARS-CoV-2 causes COVID-19, a disease curiously resulting in varied symptoms and outcomes, ranging from asymptomatic to fatal. Autoimmunity due to cross-reacting antibodies resulting from molecular mimicry between viral antigens and host proteins may provide an explanation. We computationally investigated molecular mimicry between SARS-CoV-2 Spike and known epitopes. We discovered molecular mimicry hotspots in Spike and highlight two examples with tentative autoimmune potential and implications for understanding COVID-19 complications. We show that a TQLPP motif in Spike and thrombopoietin shares similar antibody binding properties. Antibodies cross-reacting with thrombopoietin may induce thrombocytopenia, a condition observed in COVID-19 patients. Another motif, ELDKY, is shared in multiple human proteins such as PRKG1 and tropomyosin. Antibodies cross-reacting with PRKG1 and tropomyosin may cause known COVID-19 complications such as blood-clotting disorders and cardiac disease, respectively. Our findings illuminate COVID-19 pathogenesis and highlight the importance of considering autoimmune potential when developing therapeutic interventions to reduce adverse reactions.

bioinformatics

Multi-run Concrete Autoencoder to Identify Prognostic lncRNAs for 12 Cancers

Long non-coding RNA plays a vital role in changing the expression profiles of various target genes that leads to cancer development. Thus, identifying prognostic lncRNAs related to different cancers might help in developing cancer therapy. To discover the critical lncRNAs that can identify the origin of different cancers, we proposed to use the state-of-the-art deep learning algorithm Concrete Autoencoder (CAE) in an unsupervised setting, which efficiently identifies a subset of the most informative features. However, CAE does not identify reproducible features in different runs due to its stochastic nature. We proposed a multi-run CAE (mrCAE) to identify a stable set of features to address this issue. The assumption is that a feature appearing in multiple runs carries more meaningful information about the data under consideration. The genome-wide lncRNA expression profiles of 12 different types of cancers, a total of 4,768 samples available in The Cancer Genome Atlas (TCGA), were analyzed to discover the key lncRNAs. The lncRNAs identified by multiple runs of CAE were added to the final list of key lncRNAs, which are capable of identifying 12 different cancers. Our results showed that mrCAE performs better in feature selection than single-run CAE, standard autoencoder (AE), and other state-of-the-art feature selection techniques. This study discovered a set of top-ranking 128 lncRNAs that could identify the origin of 12 different cancers with an accuracy of 95%. Survival analysis showed that 76 of 128 lncRNAs have the prognostic capability in differentiating high- and low-risk groups of patients in different cancers. The proposed mrCAE outperformed AE, which select latent features and were thought to be the best tools for dimensionality reduction. By selecting actual features instead of pseudo-features, mrCAE can be valuable for precision medicine. The identified prognostic lncRNAs (this work), mRNAs, miRNAs, and DNA methylated genes (future work) can lead to biomarkers and therapies for different cancers.

bioinformatics