bioRxiv Science⌕ Search

Biology subjects

Jarwal, A.

Publications and source records attributed to Jarwal, A..

2 recordsLinked to original sources

A Deep Learning method for classification of HNSCC and HPV patients using single-cell transcriptomics

Head and Neck Squamous Cell Carcinoma (HNSCC) is the seventh most highly prevalent cancer type worldwide. Early detection of HNSCC is one of the important challenges in managing the treatment of the cancer patients. Existing techniques for detecting HNSCC are costly, expensive, and invasive in nature. In this study, we aimed to address this issue by developing classification models using machine learning and deep learning techniques, focusing on single-cell transcriptomics to distinguish between HNSCC and normal samples. Additionally, we built models to classify HNSCC samples into HPV-positive (HPV+) and HPV-negative (HPV-) categories. The models developed in this study have been trained on 80% of the GSE181919 dataset and validated on the remaining 20%. To develop an efficient model, we performed feature selection using mRMR method to shortlist a small number of genes from a plethora of genes. Artificial Neural Network based model trained on 100 genes outperformed the other classifiers with an AUROC of 0.91 for HNSCC classification for the validation set. The same algorithm achieved an AUROC of 0.83 for the classification of HPV+ and HPV-patients on the validation set. We also performed Gene Ontology (GO) enrichment analysis on the 100 shortlisted genes and found that most genes were involved in binding and catalytic activities. To facilitate the scientific community, a software package has been developed in Python which allows users to identify HNSCC in patients along with their HPV status. It is available at https://webs.iiitd.edu.in/raghava/hnscpred/. Key PointsO_LIApplication of single cell transcriptomics in cancer diagnosis C_LIO_LIDevelopment of models for predicting HNSCC patients C_LIO_LIClassification of HPV+ and HPV-HNSCC patients C_LIO_LIIdentification of gene biomarkers from single cell sequencing C_LIO_LIA standalone software package HNSCpred for predicting HNSCC patients C_LI Authors BiographyO_LIAkanksha Jarwal is currently pursuing an M. Tech. in Computational Biology at the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAnjali Dhall is currently pursuing a Ph.D. in Computational Biology at the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAkanksha Arora is currently pursuing a Ph.D. in Computational Biology at the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LISumeet Patiyal is currently pursuing a Ph.D. in Computational Biology at the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAman Srivastava is currently pursuing an M. Tech. in Computational Biology at the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIGajendra P. S. Raghava is currently working as a Professor and Head of the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LI

bioinformatics↗

Prediction of Alzheimer's Disease from Single Cell Transcriptomics Using Deep Learning

Alzheimers disease (AD) is a progressive neurological disorder characterized by brain cell death, brain atrophy, and cognitive decline. Early diagnosis of AD remains a significant challenge in effectively managing this debilitating disease. In this study, we aimed to harness the potential of single-cell transcriptomics data from 12 Alzheimers patients and 9 normal controls (NC) to develop a predictive model for identifying AD patients. The dataset comprised gene expression profiles of 33,538 genes across 169,469 cells, with 90,713 cells belonging to AD patients and 78,783 cells belonging to NC individuals. Employing machine learning and deep learning techniques, we developed prediction models. Initially, we performed data processing to identify genes expressed in most cells. These genes were then ranked based on their ability to classify AD and NC groups. Subsequently, two sets of genes, consisting of 35 and 100 genes, respectively, were used to develop machine learning-based models. Although these models demonstrated high performance on the training dataset, their performance on the validation/independent dataset was notably poor, indicating potential overoptimization. To address this challenge, we developed a deep learning method utilizing dropout regularization technique. Our deep learning approach achieved an AUC of 0.75 and 0.84 on the validation dataset using the sets of 35 and 100 genes, respectively. Furthermore, we conducted gene ontology enrichment analysis on the selected genes to elucidate their biological roles and gain insights into the underlying mechanisms of Alzheimers disease. While this study presents a prototype method for predicting AD using single-cell genomics data, it is important to note that the limited size of the dataset represents a major limitation. To facilitate the scientific community, we have created a website to provide with code and service. It is freely available at https://webs.iiitd.edu.in/raghava/alzscpred. Key PointsO_LIPredictive Model for Alzheimers Disease Using Single Cell Transcriptomics Data C_LIO_LIOveroptimization of models trained on single-cell genomics data. C_LIO_LIApplication of dropout regularization technique of ANN for reducing overoptimization C_LIO_LIRanking of genes based on their ability to predict patients Alzheimers Disease C_LIO_LIStandalone software package for predicting Alzheimers Disease C_LI Authors BiographyO_LIAman Srivastava is pursuing M. Tech. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAnjali Dhall is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LISumeet Patiyal is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAkanksha Arora is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAkanksha Jarwal is pursuing M. Tech. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIGajendra P. S. Raghava is currently working as Professor and Head of Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LI

bioinformatics↗