bioRxiv ScienceSearch

Biology subjects

Bhalla, S.

Publications and source records attributed to Bhalla, S..

4 recordsLinked to original sources

Expression based biomarkers and models to classify early and late stage samples of Papillary Thyroid Carcinoma

In this study, we describe the key transcripts and machine learning models developed for classifying the early and late stage samples of Papillary Thyroid Cancer (PTC), using transcripts expression data from The Cancer Genome Atlas (TCGA). First, we rank all the transcripts on the basis of area under receiver operating characteristic curve, (AUROC) value to discriminate the early and late stage, based on an expression threshold. With the expression of a single transcript DCN, we can classify the stage samples with a 68.5% accuracy and AUROC of 0.66. Then we implemented various combination of multiple gene panels, selected using various gold standard feature selection techniques. The model based on the expression of 36 multiple transcripts (protein coding and non-coding) selected using SVC-L1 achieves the maximum accuracy of 74.51% with AUROC of 0.75 on independent validation dataset with balanced sensitivity and specificity. Further, these signatures also performed well on external microarray data obtained from GEO, predicting nearly 70% (12 samples out of 17 samples) early stage samples correctly. Further, multiclass model, classifying the normal, early and late stage samples achieves the accuracy of 75.43% with AUROC of 0.80 on independent validation dataset. With correlation analysis, we found that transcripts with maximum change in correlation of their expression in both the stages are significantly enriched in neuroactive ligand receptor interaction pathway. We also propose a panel of five protein coding transcripts, which on the basis of their expression, can segregate cancer and normal samples with 97.32% accuracy and AUROC of 0.99 on independent validation dataset. All the models and dataset used in this study are available from the web server CancerTSP (http://webs.iiitd.edu.in/raghava/cancertsp/).

bioinformatics

Prediction and analysis of skin cancer progression using genomics profiles of patients

Metastatic state of the Skin Cutaneous Melanoma (SKCM) has led to high mortality rate worldwide. Previously, various studies have revealed the association of the metastatic melanoma with the diminished survival rate in comparison to primary tumors. Thus, prediction of melanoma at primary tumor state is crucial to employ optimal therapeutic strategy for prolonged survival of patients. The RNA, miRNA and methylation data of The Cancer Genome Atlas (TCGA) cohort of SKCM is comprehensively analysed to recognize key genomic features that can categorize various states of metastatic tumors from primary tumors with high precision. Subsequently, various prediction models were developed using filtered genomic features implementing various machine learning techniques to classify these primary tumors from metastatic tumors. The SVC model (with class weight and RBF kernel) developed using 17 mRNA features achieved maximum MCC 0.73 with sensitivity, specificity and accuracy 89.19%, 90.48% and 89.47% respectively on independent validation dataset. Our study reveals that gene expression based features performs better than features obtained from miRNA profiling and epigenomic profiling. Our analysis shows that the expression of genes C7, MMP3, KRT14, KRT17, MASP1, and miRNA hsa-mir-205 and hsa-mir-203a are among the key genomic features that may substantially contribute to the oncogenesis of melanoma even on the basis of simple expression threshold. The major prediction models and analysis modules to predict metastatic and primary tumor samples of SKCM are available from a webserver, CancerSPP (http://webs.iiitd.edu.in/raghava/cancerspp/).

bioinformatics

A web bench for analysis and prediction of oncological status from proteomics data of urine samples

Urine-based cancer biomarkers offer numerous advantages over the other biomarkers and play a crucial role in cancer management. In this study, an attempt has been made to develop proteomics-based prediction models to discriminate patients of oncological disorders related to urinary tract and healthy controls from their urine samples. The dataset used in this study was obtained from human urinary peptide database that contains urine proteomics data of 1525 oncological and 1503 healthy controls with the spectral intensity of 5605 peptides. First, we identified peptide spectra using various feature selection techniques, which display different intensity and occurrence in oncological samples and healthy controls. Based on selected 173 peptide-based biomarkers, we developed models for predicting oncological samples and achieved maximum accuracy of 91.94% with 0.84 MCC. Prediction models were also developed based on spectral intensities with known peptide sequences. We also quantitated the amount of protein in a sample based on intensities of its fragments/peptides and developed prediction models based on protein expression. It was observed that certain proteins and their peptides such as fragments of collagen protein are more abundant in oncological samples. Based on this study, we also developed a web bench, CancerUBM, for mining proteomics data, which is freely available at http://webs.iiitd.edu.in/raghava/cancerubm/.

bioinformatics

Classification of early and late stage Liver Hepatocellular Carcinoma patients from their genomics and epigenomics profiles.

BackgroundLiver Hepatocellular Carcinoma (LIHC) is the second major cancer worldwide, responsible for millions of premature deaths every year. Prediction of clinical staging is vital to implement optimal therapeutic strategy and prognostic prediction in cancer patients. However, to date, no method has been developed for predicting stage of LIHC from genomic profile of samples.\n\nResultsIn current study, in silico models have been developed for classifying LIHC patients in early and late stage using RNA expression and DNA methylation data. The Cancer Genome Atlas (TCGA) dataset contains 173 early and 177 late stage samples of LIHC, was extensively analysed to identify differentially expressed RNA transcripts and methylated CpG sites that can discriminate early and late stages of LIHC samples with high precision. Naive Bayes model developed using 51 features that combine 21 CpG methylation sites and 30 RNA transcripts achieved maximum MCC 0.58 with accuracy 78.87% on validation dataset. Further, we also analysed genomics and epigenomics profiles of normal and LIHC samples and developed model to classify LIHC samples with AUROC 0.99. In addition, multiclass models developed for classifying samples in normal, early and late stage of cancer and achieved accuracy of 76.54% and AUROC of 0.86.\n\nConclusionOur study reveals stage prediction of LIHC samples with high accuracy based on genomics and epigenomics profiling is a challenging task in comparison to classification of LIHC and normal samples. Comprehensive analysis, differentially expressed RNA transcripts, methylated CpG sites in LIHC samples and prediction models are available from CancerLSP (http://webs.iiitd.edu.in/raghava/cancerlsp/).

bioinformatics