bioRxiv Science⌕ Search

Biology subjects

Kabanga, E.

Publications and source records attributed to Kabanga, E..

3 recordsLinked to original sources

Impact of U2-type introns on splice site prediction in Arabidopsis thaliana using deep learning

In this study, we investigate the impact of introns on the effectiveness of splice site prediction using deep learning models, focusing on Arabidopsis thaliana. We specifically utilize U2-type introns due to their ubiquity in plant genomes and the rich datasets available. We formulate two hypotheses: first, that short introns would lead to a higher effectiveness of splice site prediction than long introns due to reduced spatial complexity; and second, that sequences containing multiple introns would improve prediction effectiveness by providing a richer context for splicing events. Our findings indicate that (1) models trained on datasets with shorter introns consistently outperform those trained on datasets with longer introns, highlighting the importance of intron length in splice site prediction, and (2) models trained with datasets containing multiple introns per sequence demonstrate superior effectiveness over those trained with datasets containing a single intron per sequence. Furthermore, our findings not only align with the two hypotheses we put forward but also confirm existing observations from wet lab experiments regarding the impact of length of an intron and the number of introns present in a sequence on splice site prediction effectiveness, suggesting that our computational insights come with biological relevance. Author summaryIn this study, we explore how intron characteristics affect the effectiveness of splice site predictions in Arabidopsis thaliana using deep learning. In particular, focusing on U2-type introns due to their prevalence in plant genomes and their relevance for large-scale data analysis, we demonstrate that both the length of these introns and the number of introns present in a sequence substantially influence prediction outcomes. Our findings highlight that deep learning models trained on data with shorter introns or multiple introns per sequence produce better predictions, aligning with observations from wet lab experiments regarding the impact of intron length and the number of introns per sequences on splice site prediction effectiveness.

genomics↗

Towards Interpretable Multitask Learning for Splice Site and Translation Initiation Site Prediction

In this study, we investigate the effectiveness of multi-task learning (MTL) for handling three bioinformatics tasks: donor splice site prediction, acceptor splice site prediction, and translation initiation site prediction. As the foundation for our MTL approach, we use the SpliceRover model, which has previously been successful in predicting splice sites. While providing benefits such as efficient resource utilization, reduced complexity, and streamlined model management, our findings show that the newly introduced MTL model performs comparably to the SpliceRover model trained separately for each task (single-task models), with a slight decrease in specificity, sensitivity, F1-score, and Matthews Correlation Coefficient (MCC). However, these differences are statistically insignificant (the specificity decreased with 0.0081 for acceptor splice site prediction and the MCC decreased with 0.0264 for TIS prediction), emphasizing the comparable performance of the MTL model. We further analyze the effectiveness of our MTL model using visualization techniques. The outcomes indicate that our MTL model effectively learns the relevant features associated with each task when compared to the single-task models (presence of nucleotides with a higher contribution to donor splice site prediction, polypyrimidine tracts in the upstream of acceptor splice sites, and the Kozak sequence). In conclusion, our results show that the MTL model generalizes well across all three tasks.

bioinformatics↗

Discovering Biomarker Proteins and Peptides for Parkinson's Disease Prognosis Prediction with Machine Learning and Interpretability Methods

Parkinsons disease is a neurodegenerative disorder that affects millions of people worldwide, posing significant challenges for diagnosis and treatment. This study presents a machine learning pipeline for identifying candidate biomarker proteins and peptides from cerebrospinal fluid mass spectrometry (CSF-MS) tests in Parkinsons disease patients. Our pipeline comprises two main stages: (1) model training using mutual information-based feature selection and five different machine learning regressors and (2) identification of candidate biomarkers by combining three types of interpretability methods. Our regression models demonstrated promising effectiveness in predicting the Movement Disorder Society-Unified Parkinsons Disease Rating Scale (MDS-UPDRS) scores, with UPDRS-1 receiving the best predictions, followed by UPDRS-3 and UPDRS-2. Furthermore, our pipeline identified 11 proteins and peptides as potential biomarkers for Parkinsons disease, excluding Levodopa usage which trivially has the most significant impact on the prognosis prediction. Comparisons with four additional pipelines confirmed the effectiveness of our approach in terms of both model performance and biomarker identification. In conclusion, our study presents a comprehensive machine learning pipeline that demonstrates effectiveness in predicting the severity of Parkinsons disease using CSF-MS tests. Our approach also identifies potential biomarkers, which could aid in the development of new diagnostic tools and treatments for patients with Parkinsons disease.

bioinformatics↗