bioRxiv Science⌕ Search

Biology subjects

Lamisa, A. B.

Publications and source records attributed to Lamisa, A. B..

4 recordsLinked to original sources

A Non-invasive Detection of Parkinson's Disease using PitArray: An Integrative Meta-Analysis and Machine Learning Approach

Parkinsons disease (PD) is a progressive neurodegenerative disorder affecting the central nervous system, often diagnosed in its advanced stages due to the absence of sensitive biomarkers. With this objective in mind, our study conducted a comprehensive analysis of differentially expressed genes (DEGs) sourced from blood-based microarray datasets to uncover potential biomarkers and developed a machine learning based classifier to conduct two step validations. By analyzing gene expression of three projects, we identified 678 DEGs, consisting of 337 genes showing upregulation and 341 genes presenting downregulation. Additionally, insights from functional enrichment and the protein-protein network analysis indicate that HLA-F, IRF-1, and RPS28 have the potential to serve as biomarkers for diagnosing PD. Simultaneously, we employed feature selection techniques such as Least Absolute Shrinkage and Selection Operator with Cross Validation (LassoCV) followed by Recursive Feature Elimination with Cross Validation (REFCV) to filter our initial dataset of 13,249 genes down to 43 genes, which were subsequently used to train the machine learning-based classifier models. These 43 genes formed the basis for training and testing various machine learning models, including logistic regression, random forest, naive Bayes, k-nearest neighbors, support vector machine, and deep learning based artificial neural networks. Our models demonstrated robust performance, with Support Vector Machine outperforming others by 0.65 accuracy (95%CI: 0.58-0.66), 0.70 AUC-ROC (95%CI: 0.70-0.71) and 0.35 MCC (95%CI: 0.34-0.39). The model was implemented to develop the PitArray tool for non-invasive detection of PD from blood. PitArray is available at: https://github.com/Arittra95/PitArray. Key PointsO_LIHLA-F, IRF-1, and RPS28 were identified as potential biomarkers for Parkinsons disease diagnosis. C_LIO_LISeveral sophisticated feature selection methods recognized 43 genes which were then used to build a machine learning model. C_LIO_LIA Support Vector Machine based tool named PitArray was developed which could distinguish Parkinsons disease patients from healthy people based on blood transcriptome data. C_LI

bioinformatics↗

An Integrated Comparative Genomics, Subtractive Proteomics and Immunoinformatics Framework for the Rational Design of a Pan-Salmonella Multi-Epitope Vaccine

Salmonella infections are a global public health issue due to the high cost of illness surveillance, prevention, and treatment. In this study, we explored the core proteome in Salmonella to design a multi-epitope vaccine through Subtractive Proteomics and immunoinformatics approaches. A total of 2395 core proteins presents in 30 different strains of Salmonella (reference strain-NZ CP014051) were curated. Utilizing the subtractive proteomics approach on the Salmonella core proteome, Curlin major subunit A (CsgA) was selected as the vaccine candidate. csgA is a conserved gene that is related with biofilm formation. Immunodominant B and T cell epitopes from CsgA were predicted using numerous immunoinformatics tools. T lymphocyte epitopes had adequate population coverage and their corresponding MHC alleles showed significant binding scores after peptide-protein based molecular docking. Afterward, a multiepitope vaccine was constructed with peptide linkers and Human Beta Defensin-2 (as an adjuvant). The vaccine was found to be highly antigenic, non-toxic, non-allergic, and had physicochemical properties. Additionally, Molecular Dynamics Simulation and Immune Simulation demonstrated that the vaccine can bind with Toll Like Receptor 4 and elicit robust immune response. Using in vitro, in vivo, and clinical trials, our results would yield a Pan-Salmonella vaccine that will provide protection against various Salmonella species.

bioinformatics↗

A meta-analysis of bulk RNA-seq datasets identifies potential biomarkers and repurposable therapeutics against Alzheimers disease

Alzheimers disease (AD) poses a major challenge due to its impact on the elderly population and the lack of effective early diagnosis and treatment options. In an effort to address this issue, a study focused on identifying potential biomarkers and therapeutic agents for AD was carried out. Using RNA-Seq data from AD patients and healthy individuals, 12 differentially expressed genes (DEGs) were identified, with 9 expressing upregulation (ISG15, HRNR, MTATP8P1, MTCO3P12, DTHD1, DCX, ST8SIA2, NNAT, and PCDH11Y) and 3 expressing downregulation (LTF, XIST, and TTR). Among them, TTR exhibited the lowest gene expression profile. Interestingly, functional analysis tied TTR to amyloid fiber formation and neutrophil degranulation through enrichment analysis. These findings suggested the potential of TTR as a diagnostic biomarker for AD. Additionally, druggability analysis revealed that the FDA-approved drug Levothyroxine might be effective against the Transthyretin protein encoded by the TTR gene. Molecular docking and dynamics simulation studies of Levothyroxine and Transthyretin suggested that this drug could be repurposed to treat AD. However, additional studies using in vitro and in vivo models are necessary before these findings can be applied in clinical applications.

bioinformatics↗

AITeQ: A machine learning framework for Alzheimer's prediction using a distinctive 5-gene signature

Neurodegenerative diseases, such as Alzheimers disease, pose a significant global health challenge with their complex etiology and elusive biomarkers. In this study, we developed the Alzheimers Identification Tool using RNA-Seq (AITeQ), a machine learning model based on an optimized random forest algorithm for identification of Alzheimers from RNA-Seq data. Analysis of RNA-Seq data from 433 individuals, including 293 Alzheimers patients and 140 controls led to the discovery of 47,929 differentially expressed genes. This was followed by a machine learning protocol involving feature selection, model training, performance evaluation, and hyperparameter tuning. The feature selection process undertaken in this study, employing a combination of 4 different methodologies, culminated in the identification of a compact yet impactful set of 5 genes. Ten diverse machine learning models were trained and tested using these 5 genes (ITGA10, CXCR4, ADCYAP1, SLC6A12, VGF). Performance metrics, including precision, recall, F1-score, accuracy, receiver operating characteristic area under the curve, and confusion matrices, were assessed before and after hyperparameter tuning. Overall, the random forest model with optimized hyperparameters was identified as the best and was used to develop AITeQ. AITeQ is available at: https://github.com/ishtiaque-ahammad/AITeQ Key PointsO_LIA set of 5 genes (ITGA10, CXCR4, ADCYAP1, SLC6A12, VGF) were identified following differential gene expression and feature importance analysis. C_LIO_LITen diverse machine learning algorithms were trained and tested using the gene expression patterns of the identified 5 genes. The random forest algorithm with customized hyperparameters was found to be the best-performing model for differentiating Alzheimers disease samples from control. C_LIO_LIAITeQ, a user-friendly, reliable, and accurate machine learning framework for Alzheimers disease prediction was developed based on the 5 gene signature. C_LI

bioinformatics↗