bioRxiv Science⌕ Search

Biology subjects

Enoma, D. O.

Publications and source records attributed to Enoma, D. O..

2 recordsLinked to original sources

Extreme gradient boosting machine learning algorithm identifies genome-wide relevant genetic variants in prostate cancer risk prediction.

Genome-wide association studies (GWAS) identify the variants (Single Nucleotide polymorphisms) associated with a disease phenotype within populations. These genetic differences are essential in variations in incidence and mortalities, especially for Prostate cancer in the African population. Given the complexity of cancer, it is imperative to identify the variants that contribute to the development of the disease. The standard univariate analysis employed in GWAS may not capture the non-linear additive interactions between variants, which might affect the risk of developing Prostate cancer. This is because the interactions in complex diseases such as prostate cancer are usually non-linear and would benefit from a non-linear Machine Learning gradient boosting viz XGBoost (extreme gradient boosting). We applied the XGBoost algorithm and an iterative SNP selection algorithm to find the top features (SNPs) that best predict the risk of developing prostate cancer with a Support Vector Machine (SVM). The number of subjects was 907, and input features were 1,798,727 after appropriate quality control. The algorithm involved ten trials of 5-fold cross-validation to optimize the datasets hyperparameters and the prediction tasks second module (utilizing SVM). The model achieved AUC-ROC cure of 0.66, 0.57 and 0.55 on the Train, Dev and Test sets, respectively. The area under the Precision-Recall Curve was 0.69, 0.60 and 0.57 on the Train, Dev and Test sets, respectively. Furthermore, the final number of predictive risk variants was 2798, associated with 847 Ensembl genes. Interaction analysis showed that Nodes were 339 and the edges were 622 in the gene interaction network. This shows evidence that the non-linear Machine learning approach offers excellent possibilities for understanding the genetic basis of complex diseases.

bioinformatics↗

Multi-omics data integration approach identifies potential biomarkers for Prostate cancer

Prostate cancer (PCa) is one of the most common malignancies, and many studies have shown that PCa has a poor prognosis, which varies across different ethnicities. This variability is caused by genetic diversity. High-throughput omics technologies have identified and shed some light on the mechanisms of its progression and finding new biomarkers. Still, a systems biology approach is needed for a holistic molecular perspective. In this study, we applied a multi-omics approach to data analysis using different publicly available omics data sets from diverse populations to better understand the PCa disease etiology. Our study used multiple omic datasets, which included genomic, transcriptomic and metabolomic datasets, to identify drivers for PCa better. Individual omics datasets were analysed separately based on the standard pipeline for each dataset. Furthermore, we applied a novel multi-omics pathways algorithm to integrate all the individual omics datasets. This algorithm applies the p-values of enriched pathways from unique omics data types, which are then combined using the MiniMax statistic of the PathwayMultiomics tool to prioritise pathways dysregulated in the omics datasets. The single omics result indicated an association between up-regulated genes in RNA-Seq data and the metabolomics data. Glucose and pyruvate are the primary metabolites, and the associated pathways are glycolysis, gluconeogenesis, pyruvate kinase deficiency, and the Warburg effect pathway. From the interim result, the identified genes in RNA-Seq single omics analysis are linked with the significant pathways from the metabolomics analysis. The multi-omics pathway analysis will eventually enable the identification of biomarkers shared amongst these different omics datasets to ease prostate cancer prognosis.

bioinformatics↗