bioRxiv Science⌕ Search

Biology subjects

Bajiya, N.

Publications and source records attributed to Bajiya, N..

5 recordsLinked to original sources

Prediction of plant resistance proteins using alignment-based and alignment-free approaches

Plant Disease Resistance (PDR) proteins are critical in identifying and killing plant pathogens. Predicting PDR protein is essential for understanding plant-pathogen interactions and developing strategies for crop protection. This study proposes a hybrid model for predicting and designing PDR proteins against plant-invading pathogens. Initially, we tried alignment-based approaches, such as BLAST for similarity search and MERCI for motif search. These alignment-based approaches exhibit very poor coverage or sensitivity. To overcome these limitations, we developed alignment-free or machine learning-based methods using compositional features of proteins. Our machine learning-based model, developed using compositional features of proteins, achieved a maximum performance AUROC of 0.92. The performance of our model improved significantly from AUROC of 0.92 to 0.95 when we used evolutionary information instead of protein sequence. Finally, we developed a hybrid or ensemble model that combined our best machine learning model with BLAST and obtained the highest AUROC of 0.98 on the validation dataset. We trained and tested our models on a training dataset and evaluated them on a validation dataset. None of the proteins in our validation dataset are more than 40% similar to proteins in the training dataset. One of the objectives of this study is to facilitate the scientific community working in plant biology. Thus, we developed an online platform for predicting and designing plant resistance proteins, "PlantDRPpred" (https://webs.iiitd.edu.in/raghava/plantdrppred). HighlightsO_LIDevelopment of a Machine-learning model for resistance protein prediction. C_LIO_LIUsed alignment-based and alignment-free ensemble methods. C_LIO_LIWeb server development and standalone package. C_LIO_LIPrediction and design of PDR proteins. C_LI

bioinformatics↗

Prediction of inhibitory peptides against E. coli with desired MIC value

In the past, several methods have been developed for predicting antibacterial and antimicrobial peptides, but only limited attempts have been made to predict their minimum inhibitory concentration (MIC) values. In this study, we trained our models on 3,143 peptides and validated them on 786 peptides whose MIC values have been determined experimentally against Escherichia coli (E. coli). The correlational analysis reveals that the Composition Enhanced Transition and Distribution (CeTD) attributes strongly correlate with MIC values. We initially employed the similarity search strategy utilizing BLAST to estimate MIC values of peptides but found it inadequate for prediction. Next, we developed machine learning techniques-based regression models using a wide range of features, including peptide composition, binary profile, and embeddings of large language models. We implemented feature selection techniques like minimum Redundancy Maximum Relevance (mRMR) to select the best relevant features for developing prediction models. Our Random forest-based regressor, based on selected features, achieved a correlation coefficient (R) of 0.78, R-squared (R{superscript 2}) of 0.59, and a root mean squared error (RMSE) of 0.53 on the validation dataset. Our best model outperforms the existing methods when benchmarked on an independent dataset of 498 inhibitory peptides of E. coli. One of the major features of the web-based platform EIPpred developed in this study is that it allows users to identify or design peptides that can inhibit E. coli with the desired MIC value (https://webs.iiitd.edu.in/raghava/eippred). HighlightsO_LIPrediction of MIC value of peptides against E.coli. C_LIO_LIAn independent dataset was generated for comparison. C_LIO_LIFeature selection using the mRMR method. C_LIO_LIA regressor method for designing novel inhibitory peptides. C_LIO_LIA web server and standalone package for predicting the inhibitory activity of peptides. C_LI

bioinformatics↗

Prediction of anti-freezing proteins from their evolutionary profile

1.Prediction of antifreeze proteins (AFPs) holds significant importance due to their diverse applications in healthcare. An inherent limitation of current AFP prediction methods is their reliance on unreviewed proteins for evaluation. This study evaluates proposed and existing methods on an independent dataset containing 81 AFPs and 73 non-AFPs obtained from Uniport, which have been already reviewed by experts. Initially, we constructed machine learning models for AFP prediction using selected composition-based protein features and achieved a peak AUC of 0.90 with an MCC of 0.69 on the independent dataset. Subsequently, we observed a notable enhancement in model performance, with the AUC increasing from 0.90 to 0.93 upon incorporating evolutionary information instead of relying solely on the primary sequence of proteins. Furthermore, we explored hybrid models integrating our machine learning approaches with BLAST-based similarity and motif-based methods. However, the performance of these hybrid models either matched or was inferior to that of our best machine-learning model. Our best model based on evolutionary information outperforms all existing methods on independent/validation dataset. To facilitate users, a user-friendly web server with a standalone package named "AFPropred" was developed (https://webs.iiitd.edu.in/raghava/afpropred). HighlightsO_LIPrediction of antifreeze proteins with high precision C_LIO_LIEvaluation of prediction models on an independent dataset C_LIO_LIMachine learning based models using sequence composition C_LIO_LIEvolutionary information based prediction models C_LIO_LIA webserver for predicting, scanning, and designing AFPs. C_LI Authors BiographyO_LINishant Kumar is currently working as Ph.D. in Computational biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIShubham Choudhury is currently working as Ph.D. in Computational biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India C_LIO_LINisha Bajiya is currently working as Ph.D. in Computational biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India C_LIO_LISumeet Patiyal is currently working as a postdoctoral visiting fellow Cancer Data Science Laboratory, National Cancer Institute, National Institutes of Health, Bethesda, Maryland, USA. C_LIO_LIGajendra P. S. Raghava is currently working as Professor and Head of Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LI

bioinformatics↗

AntiBP3: A hybrid method for predicting antibacterial peptides against gram-positive/negative/variable bacteria

This study focuses on the development of in silico models for predicting antibacterial peptides as a potential solution for combating antibiotic-resistant strains of bacteria. Existing methods for predicting antibacterial peptides are mostly designed to target either gram-positive or gram-negative bacteria. In this study, we introduce a novel approach that enables the prediction of antibacterial peptides against several bacterial groups, including gram-positive, gram-negative, and gram-variable bacteria. Firstly, we developed an alignment-based approach using BLAST to identify antibacterial peptides and achieved poor sensitivity. Secondly, we employed a motif-based approach to predict antibacterial peptides and obtained high precision with low sensitivity. To address the similarity issue, we developed machine learning-based models using a variety of compositional and binary features. Our machine learning-based model developed using the amino acid binary profile of terminal residues achieved maximum AUC 0.93, 0.98 and 0.94 for gram-positive, gram-negative, and gram-variable bacteria, respectively, when evaluated on a validation/independent dataset. Our attempts to develop hybrid or ensemble methods by merging machine learning models with similarity and motif-based techniques did not yield any improvements. To ensure robust evaluation, we employed standard techniques such as five-fold cross-validation, internal validation, and external validation. Our method performs better than existing methods when we compare our method with existing approaches on an independent dataset. In summary, this study makes significant contributions to the field of antibacterial peptide prediction by providing a comprehensive set of methods tailored to different bacterial groups. As part of our contribution, we have developed the AntiBP3 web server and standalone package, which will assist researchers in the discovery of novel antibacterial peptides for combating bacterial infections (https://webs.iiitd.edu.in/raghava/antibp3/). Key Points BLAST-based similarity for annotating antibacterial peptides. Machine learning-based models developed using composition and binary profiles. Identification and mapping of motifs exclusively found in antibacterial peptides Improved version of AntiBP and AntiBP2 for predicting antibacterial peptides. Web server for predicting/designing/scanning antibacterial peptides for all groups of bacteria Authors BiographyO_LINisha Bajiya is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIShubham Choudhury is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAnjali Dhall is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIGajendra P. S. Raghava is currently working as Professor and Head of Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LI

bioinformatics↗

A hybrid approach for predicting multi-label subcellular localization of mRNA at genome scale

In the past, number of methods have been developed for predicting single label subcellular localization of mRNA in a cell. Only limited methods had been built to predict multi-label subcellular localization of mRNA. Most of the existing methods are slow and cannot be implemented at transcriptome scale. In this study, a fast and reliable method had been developed for predicting multi-label subcellular localization of mRNA that can be implemented at genome scale. Firstly, deep learning method based on convolutional neural network method have been developed using one-hot encoding and attained an average AUROC - 0.584 (0.543 - 0.605). Secondly, machine learning based methods have been developed using mRNA sequence composition, our XGBoost classifier achieved an average AUROC - 0.709 (0.668 - 0.732). In addition to alignment free methods, we also developed alignment-based methods using similarity and motif search techniques. Finally, a hybrid technique has been developed that combine XGBoost models and motif-based searching and achieved an average AUROC 0.742 (0.708 - 0.816). Our method - MRSLpred, developed in this study is complementary to the existing method. One of the major advantages of our method over existing methods is its speed, it can scan all mRNA of a transcriptome in few hours. A publicly accessible webserver and a standalone tool has been developed to facilitate researchers (Webserver: https://webs.iiitd.edu.in/raghava/mrslpred/). Key PointsO_LIPrediction of Subcellular localization of mRNA C_LIO_LIClassification of mRNA based on Motif and BLAST search C_LIO_LICombination of alignment based and alignment free techniques C_LIO_LIA fast method for subcellular localization of mRNA C_LIO_LIA web server and standalone software C_LI

bioinformatics↗