bioRxiv Science⌕ Search

Biology subjects

Noor, H.

Publications and source records attributed to Noor, H..

2 recordsLinked to original sources

Revealing cancer driver genes through integrative transcriptomic and epigenomic analyses with Moonlight

Cancer involves dynamic changes caused by (epi)genetic alterations such as mutations or abnormal DNA methylation patterns which occur in cancer driver genes. These driver genes are divided into oncogenes and tumor suppressors depending on their function and mechanism of action. Discovering driver genes in different cancer (sub)types is important not only for increasing current understanding of carcinogenesis but also from prognostic and therapeutic perspectives. We have previously developed a framework called Moonlight which uses a systems biology multi-omics approach for prediction of driver genes. Here, we present an important development in Moonlight2 by incorporating a DNA methylation layer which provides epigenetic evidence for deregulated expression profiles of driver genes. To this end, we present a novel functionality called Gene Methylation Analysis (GMA) which investigates abnormal DNA methylation patterns to predict driver genes. This is achieved by integrating the tool EpiMix which is designed to detect such aberrant DNA methylation patterns in a cohort of patients and further couples these patterns with gene expression changes. To showcase GMA, we applied it to three cancer (sub)types (basal-like breast cancer, lung adenocarcinoma, and thyroid carcinoma) where we discovered 33, 190, and 263 epigenetically driven genes, respectively. A subset of these driver genes had prognostic effects with expression levels significantly affecting survival of the patients. Moreover, a subset of the driver genes demonstrated therapeutic potential as drug targets. This study provides a framework for exploring the driving forces behind cancer and provides novel insights into the landscape of three cancer sub(types) by integrating gene expression and methylation data. Moonlight2R is available on GitHub (https://github.com/ELELAB/Moonlight2R) and BioCondcutor (https://bioconduc-tor.org/packages/release/bioc/html/Moonlight2R.html). The associated case studies presented here are available on GitHub (https://github.com/ELELAB/Moon-light2_GMA_case_studies) and OSF (https://osf.io/j4n8q/). Author summaryCancer is a complex disease and a main cause of mortality worldwide. This heterogeneous disease arises due to accumulation of changes which occur in driver genes that drive cancer progression when they are altered. These driver genes are commonly divided into oncogenes, which promote cancer, and tumor suppressors, which prevent it. A major goal of cancer research is identifying these driver genes, crucial for increasing our current understanding of cancer biology and for developing novel treatment approaches. A large number of cancer driver genes have already been identified. However, the underlying mechanisms for the alterations in these genes is challenging to predict given their context-dependent behavior and the complexity of cancer. Such explanations are the focus of this study with the aim of providing evidence of why certain genes do not function normally in cancer. Within this context, we present new functionalities to our previously developed cancer driver predictive framework, Moonlight. These new functionalities integrate multiple data types to predict oncogenes and tumor suppressors in a systems-biology-oriented manner that is freely available as a R package for the community.

bioinformatics↗

Digital profiling of cancer transcriptomes from histology images with grouped vision attention

Cancer is a heterogeneous disease that demands precise molecular profiling for better understanding and management. Recently, deep learning has demonstrated potentials for cost-efficient prediction of molecular alterations from histology images. While transformer-based deep learning architectures have enabled significant progress in non-medical domains, their application to histology images remains limited due to small dataset sizes coupled with the explosion of trainable parameters. Here, we develop SEQUOIA, a transformer model to predict cancer transcriptomes from whole-slide histology images. To enable the full potential of transformers, we first pre-train the model using data from 1,802 normal tissues. Then, we fine-tune and evaluate the model in 4,331 tumor samples across nine cancer types. The prediction performance is assessed at individual gene levels and pathway levels through Pearson correlation analysis and root mean square error. The generalization capacity is validated across two independent cohorts comprising 1,305 tumors. In predicting the expression levels of 25,749 genes, the highest performance is observed in cancers from breast, kidney and lung, where SEQUOIA accurately predicts the expression of 11,069, 10,086 and 8,759 genes, respectively. The accurately predicted genes are associated with the regulation of inflammatory response, cell cycles and metabolisms. While the model is trained at the tissue level, we showcase its potential in predicting spatial gene expression patterns using spatial transcriptomics datasets. Leveraging the prediction performance, we develop a digital gene expression signature that predicts the risk of recurrence in breast cancer. SEQUOIA deciphers clinically relevant gene expression patterns from histology images, opening avenues for improved cancer management and personalized therapies.

bioinformatics↗