bioRxiv · 10.1101/2025.02.12.637996
StackGlyEmbed: Prediction of N-linked Glycosylation sites using protein language models
Abstract
N-linked glycosylation is one of the most basic post-translational modifications (PTMs) where oligosaccharides covalently bond with Asparagine (N). These are found in the conserved regions like N-X-S or N-X-T where X can be any residue except Proline (P). Prediction of N-linked glycosylation sites has great importance as these PTMs play a vital role in many biological processes and functionalities. Experimental methods, such as mass spectrometry, for detecting N-linked glycosylation sites are very expensive. Therefore, prediction of N-linked glycosylation sites has become an important research field. In this work, we propose StackGlyEmbed, a stacking ensemble machine learning model, to computationally predict N-linked glycosylation sites. We have explored embeddings from several protein language models and built the stacking ensemble using SVM, XGB and KNN learners in the base layer, with a second SVM model in the meta layer. StackGlyEmbed achieves 98.2% sensitivity, 92.5% balanced accuracy, 89.1% F1-score and 82.6% MCC in independent testing, outperforming the existing SOTA methods. StackGlyEmbed is freely available at https://github.com/nafcoder/StackGlyEmbed.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nafi, M. M. I., Rahman, M. S.. 2025-02-14. StackGlyEmbed: Prediction of N-linked Glycosylation sites using protein language models. https://doi.org/10.1101/2025.02.12.637996
Cite the original work for its findings. Save a collection to share your selection of sources.