bioRxiv ScienceSearch

Biology subjects

Raghava, G.

Publications and source records attributed to Raghava, G..

2 recordsLinked to original sources

Expression based biomarkers and models to classify early and late stage samples of Papillary Thyroid Carcinoma

In this study, we describe the key transcripts and machine learning models developed for classifying the early and late stage samples of Papillary Thyroid Cancer (PTC), using transcripts expression data from The Cancer Genome Atlas (TCGA). First, we rank all the transcripts on the basis of area under receiver operating characteristic curve, (AUROC) value to discriminate the early and late stage, based on an expression threshold. With the expression of a single transcript DCN, we can classify the stage samples with a 68.5% accuracy and AUROC of 0.66. Then we implemented various combination of multiple gene panels, selected using various gold standard feature selection techniques. The model based on the expression of 36 multiple transcripts (protein coding and non-coding) selected using SVC-L1 achieves the maximum accuracy of 74.51% with AUROC of 0.75 on independent validation dataset with balanced sensitivity and specificity. Further, these signatures also performed well on external microarray data obtained from GEO, predicting nearly 70% (12 samples out of 17 samples) early stage samples correctly. Further, multiclass model, classifying the normal, early and late stage samples achieves the accuracy of 75.43% with AUROC of 0.80 on independent validation dataset. With correlation analysis, we found that transcripts with maximum change in correlation of their expression in both the stages are significantly enriched in neuroactive ligand receptor interaction pathway. We also propose a panel of five protein coding transcripts, which on the basis of their expression, can segregate cancer and normal samples with 97.32% accuracy and AUROC of 0.99 on independent validation dataset. All the models and dataset used in this study are available from the web server CancerTSP (http://webs.iiitd.edu.in/raghava/cancertsp/).

bioinformatics

A web bench for analysis and prediction of oncological status from proteomics data of urine samples

Urine-based cancer biomarkers offer numerous advantages over the other biomarkers and play a crucial role in cancer management. In this study, an attempt has been made to develop proteomics-based prediction models to discriminate patients of oncological disorders related to urinary tract and healthy controls from their urine samples. The dataset used in this study was obtained from human urinary peptide database that contains urine proteomics data of 1525 oncological and 1503 healthy controls with the spectral intensity of 5605 peptides. First, we identified peptide spectra using various feature selection techniques, which display different intensity and occurrence in oncological samples and healthy controls. Based on selected 173 peptide-based biomarkers, we developed models for predicting oncological samples and achieved maximum accuracy of 91.94% with 0.84 MCC. Prediction models were also developed based on spectral intensities with known peptide sequences. We also quantitated the amount of protein in a sample based on intensities of its fragments/peptides and developed prediction models based on protein expression. It was observed that certain proteins and their peptides such as fragments of collagen protein are more abundant in oncological samples. Based on this study, we also developed a web bench, CancerUBM, for mining proteomics data, which is freely available at http://webs.iiitd.edu.in/raghava/cancerubm/.

bioinformatics