bioRxiv Science⌕ Search

Biology subjects

Abeln, S.

Publications and source records attributed to Abeln, S..

3 recordsLinked to original sources

Proteome encoded determinants of protein sorting into extracellular vesicles

Extracellular vesicles (EVs) are membranous structures released by cells into the extracellular space and are thought to be involved in cell-to-cell communication. While EVs and their cargo are promising biomarker candidates, protein sorting mechanisms of proteins to EVs remain unclear. In this study, we ask if it is possible to determine EV association based on the protein sequence. Additionally, we ask what the most important determinants are for EV association. We answer these questions with explainable AI models, using human proteome data from EV databases to train and validate the model. It is essential to correct the datasets for contaminants introduced by coarse EV isolation workflows and for experimental bias caused by mass spectrometry. In this study, we show that it is indeed possible to predict EV association from the protein sequence: a simple sequence-based model for predicting EV proteins achieved an area under the curve of 0.77{+/-}0.01, which increased further to 0.84{+/-}0.00 when incorporating curated post-translational modification (PTM) annotations. Feature analysis shows that EV associated proteins are stable, polar, and structured with low isoelectric point compared to non-EV proteins. PTM annotations emerged as the most important features for correct classification; specifically palmitoylation is one of the most prevalent EV sorting mechanisms for unique proteins. Palmitoylation and nitrosylation sites are especially prevalent in EV proteins that are determined by very strict isolation protocols, indicating they could potentially serve as quality control criteria for future studies. This computational study offers an effective sequence-based predictor of EV associated proteins with extensive characterisation of the human EV proteome that can explain for individual proteins which factors contribute to their EV association.

bioinformatics↗

PIPENN: Protein Interface Prediction with an Ensemble of Neural Nets

MotivationProtein interactions play an essential role in many biological and cellular processes, such as protein-protein interaction (PPI) in signaling pathways, binding to DNA in transcription, and binding to small molecules in receptor activation or enzymatic activity. Experimental identification of protein binding interface residues is a time-consuming, costly, and challenging task. Several machine learning and other computational approaches exist which predict such interface residues. Here we explore if Deep Learning (DL) can be used effectively for this prediction task, and which learning strategies and architectures may be most efficient. We introduce seven DL architectures that are applied to eleven independent test sets, focused on the residues involved in PPI interfaces and in binding RNA/DNA and small molecule ligands. ResultsWe constructed a large data set dubbed BioDL, comprising protein-protein interaction data from the PDB and protein-ligand interactions (DNA, RNA and small molecules) from the BioLip database. Additionally, we reused our existing curated homo- and heteromeric PPI data sets. We performed several experiments to assess the impact of different data features, spatial forms, encoding schemes, network initializations, loss functions, regularization mechanisms, and activation functions on the performance of the predictors. Benchmarking the resulting DL models with an independent test set (ZK448) shows no single DL architecture performs best on all instances, but that an ensemble of DL architectures consistently achieves peak prediction performance. Our PIPENNs ensemble predictor outperforms current state-of-the-art sequence-based protein interface predictors on all interaction types, achieving AUCs of 0.718 (protein-protein), 0.823 (protein-nucleotide) and 0.842 (protein- small molecule) respectively. AvailabilitySource code and data sets at https://github.com/ibivu/pipenn/ Contactr.haydarlou@vu.nl

bioinformatics↗

SeRenDIP-CE: Sequence-based Interface Prediction for Conformational Epitopes

MotivationAntibodies play an important role in clinical research and biotechnology, with their specificity determined by the interaction with the antigens epitope region, as a special type of protein-protein interaction (PPI) interface. The ubiquitous availability of sequence data, allows us to predicting epitopes from sequence in order to focus time-consuming wet-lab experiments onto the most promising epitope regions. Here, we extend our previously developed sequence-based predictors for homodimer and heterodimer PPI interfaces to predict epitope residues that have the potential to bind an antibody. ResultsWe collected and curated a high quality epitope dataset from the SAbDaB database. Our generic PPI heterodimer predictor obtained an AUC-ROC of 0.666 when evaluated on the epitope test set. We then trained a random forest model specifically on the epitope dataset, reaching AUC 0.694. Further training on the combined heterodimer and epitope datasets, improves our final predictor to AUC 0.703 on the epitope test set. This is better than the best state-of-the-art sequence-based epitope predictor BepiPred-2.0. On one solved antibody-antigen structure of the COVID19 virus spike RNA binding domain, our predictor reaches AUC 0.778. We added the SeRenDIP-CE Conformational Epitope predictors to our webserver, which is simple to use and only requires a single antigen sequence as input, which will help make the method immediately applicable in a wide range of biomedical and biomolecular research. AvailabilityWebserver, source code and datasets are available at www.ibi.vu.nl/programs/serendipwww/ Contactk.a.feenstra@vu.nl

bioinformatics↗