bioRxiv Science⌕ Search

Biology subjects

Feenstra, K. A. A.

Publications and source records attributed to Feenstra, K. A. A..

2 recordsLinked to original sources

PIPENN: Protein Interface Prediction with an Ensemble of Neural Nets

MotivationProtein interactions play an essential role in many biological and cellular processes, such as protein-protein interaction (PPI) in signaling pathways, binding to DNA in transcription, and binding to small molecules in receptor activation or enzymatic activity. Experimental identification of protein binding interface residues is a time-consuming, costly, and challenging task. Several machine learning and other computational approaches exist which predict such interface residues. Here we explore if Deep Learning (DL) can be used effectively for this prediction task, and which learning strategies and architectures may be most efficient. We introduce seven DL architectures that are applied to eleven independent test sets, focused on the residues involved in PPI interfaces and in binding RNA/DNA and small molecule ligands. ResultsWe constructed a large data set dubbed BioDL, comprising protein-protein interaction data from the PDB and protein-ligand interactions (DNA, RNA and small molecules) from the BioLip database. Additionally, we reused our existing curated homo- and heteromeric PPI data sets. We performed several experiments to assess the impact of different data features, spatial forms, encoding schemes, network initializations, loss functions, regularization mechanisms, and activation functions on the performance of the predictors. Benchmarking the resulting DL models with an independent test set (ZK448) shows no single DL architecture performs best on all instances, but that an ensemble of DL architectures consistently achieves peak prediction performance. Our PIPENNs ensemble predictor outperforms current state-of-the-art sequence-based protein interface predictors on all interaction types, achieving AUCs of 0.718 (protein-protein), 0.823 (protein-nucleotide) and 0.842 (protein- small molecule) respectively. AvailabilitySource code and data sets at https://github.com/ibivu/pipenn/ Contactr.haydarlou@vu.nl

bioinformatics↗

Online biophysical predictions for SARS-CoV-2 proteins

The SARS-CoV-2 virus, the causative agent of COVID-19, consists of an assembly of proteins that determine its infectious and immunological behavior, as well as its response to therapeutics. Major structural biology efforts on these proteins have already provided essential insights into the mode of action of the virus, as well as avenues for structure-based drug design. However, not all of the SARS-CoV-2 proteins, or regions thereof, have a well-defined three-dimensional structure, and as such might exhibit ambiguous, dynamic behaviour that is not evident from static structure representations, nor from molecular dynamics simulations using these structures. We here present a website (http://sars2.bio2byte.be/) that provides protein sequence-based predictions of the backbone and side-chain dynamics and conformational propensities of these proteins, as well as derived early folding, disorder, {beta}-sheet aggregation and protein-protein interaction propensities. These predictions attempt to capture the emergent properties of the proteins, so the inherent biophysical propensities encoded in the sequence, rather than context-dependent behaviour such as the final folded state. In addition, we provide an indication of the biophysical variation that is observed in homologous proteins, which give an indication of the limits of the functionally relevant biophysical behaviour of these proteins. With this website, we therefore hope to provide researchers with further clues on the behaviour of SARS-CoV-2 proteins.

biophysics↗