bioRxiv Science⌕ Search

Biology subjects

Stanev, V.

Publications and source records attributed to Stanev, V..

2 recordsLinked to original sources

ACCURATE PREDICTION OF ASPARAGINE DEAMIDATION IN BIOLOGICS USING ADVANCED MACHINE LEARNING MODELS

The spontaneous deamidation of asparagine residues remains a major obstacle to the stability and efficacy of protein therapeutics. Currently available models in the literature for predicting deamidation liabilities can suffer from limited generalizability, likely due to biases such as sequence similarity within datasets. In this study, we built machine learning models using protein language models (e.g., ESM2) and graph neural networks (GNNs), trained on a comprehensive dataset of 591 asparagine sites from over 105 protein molecules. To address the critical issue of data leakage, we implemented a peptide grouping strategy yielding more accurate estimates of model performance for novel deamidation sites. Our analysis shows that, when sequence similarity bias is controlled, protein language models match traditional feature-based models that use amino acid composition, k-mers, PSSMs, and predicted secondary structure/solvent accessibility, while offering substantial computational advantages. Additionally, our GNN-based pipeline further increases prediction accuracy by up to 8% compared to language model-only tools and delivers a 15-25% improvement over motif-based approaches. This methodological framework enables more reliable and rapid in-silico prediction of deamidation liabilities, potentially reducing costly late-stage interventions in protein therapeutic development and is generalizable to the modeling of additional protein post-translational modifications.

bioengineering↗

Accelerating Antibody Development: Sequence and Structure-Based Models for Predicting Developability Properties through Size Exclusion Chromatography

Experimental screening for biopharmaceutical developability properties typically relies on resource-intensive, and time-consuming assays such as size exclusion chromatography (SEC). This study highlights the potential of in silico models to accelerate the screening process by exploring sequence and structure-based machine learning techniques. Specifically, we compared surrogate models based on pre-computed features extracted from sequence and predicted structure with sequence-based approaches using protein language models (PLMs) like ESM-2. In addition to different end-to-end fine-tuning strategies for PLM, we have also investigated the integration of the structural information of the antibodies into the prediction pipeline through graph neural networks (GNN). We applied these different methods for predicting protein aggregation propensity using a dataset of approximately 1200 Immunoglobulin G (IgG1) molecules. Through this empirical evaluation, our study identifies the most effective in silico approach for predicting developability properties for SEC assays, thereby adding insights to existing screening efforts for accelerating the antibody development process.

bioinformatics↗