bioRxiv ScienceSearch

Biology subjects

Staritzbichler, R.

Publications and source records attributed to Staritzbichler, R..

2 recordsLinked to original sources

ProteinPrompt: a webserver for predictingprotein-protein interactions

MotivationProtein-protein interactions play an essential role in a great variety of cellular processes and are therefore of significant interest for the design of new therapeutic compounds as well as the identification of side-effects due to unexpected binding. Here, we present ProteinPrompt, a webserver that uses machine-learning algorithms to calculate specific, currently unknown protein-protein interactions. Our tool is designed to quickly and reliably predict contacts based on an input sequence in order to scan large sequence libraries for potential binding partners, with the goal to accelerate and assure the quality of the laborious process of drug target identification. MethodsWe collected and thoroughly filtered a comprehensive database of known contacts from several sources, which is available as download. ProteinPrompt provides two complementary search methods of similar accuracy for comparison and consensus building. The default method is a random forest algorithm that uses the auto-correlations of seven amino acid scales. Alternatively, a graph neural network implementation can be selected. Additionally, a consensus prediction is available. For each query sequence, potential binding partners are identified from a protein sequence database. The proteom of several organisms are available and can be searched for contacts. ResultsTo evaluate the predictive power of the algorithms, we prepared a test dataset that was rigorously filtered for redundancy. No sequence pairs similar to the ones used for training were included in this dataset. With this challenging dataset, the random forest method achieved an accuracy rate of 0.88 and an area under curve of 0.95. The graph neural network achieved an accuracy rate of 0.86 using the same dataset. Since the underlying learning approaches are unrelated, comparing the results of random forest and graph neural networks reduces the likelihood of errors. The consensus reached an accuracy of 0.89. ProteinPrompt is available online at: http://proteinformatics.org/ProteinPrompt The server makes it possible to scan the human proteome for potential binding partners of an input sequence within minutes. For local offline usage, we furthermore created a ProteinPrompt Docker image which allows for batch submission: https://gitlab.hzdr.de/Proteinprompt/ProteinPrompt. In conclusion, we offer a fast, accurate, easy-to-use online service for predicting binding partners from an input sequence.

bioinformatics

Refining pairwise sequence alignments of membrane proteins by the incorporation of anchors

The alignment of primary sequences is a fundamental step in the analysis of protein structure, function, and evolution. Integral membrane proteins pose a significant challenge for such sequence alignment approaches, because their evolutionary relationships can be very remote, and because a high content of hydrophobic amino acids reduces their complexity. Frequently, biochemical or biophysical data is available that informs the optimum alignment, for example, indicating specific positions that share common functional or structural roles. Currently, if those positions are not correctly aligned by a standard pairwise alignment procedure, the incorporation of such information into the alignment is typically addressed in an ad hoc manner, with manual adjustments. However, such modifications are problematic because they reduce the robustness and reproducibility of the alignment. An alternative approach is the use of restraints, or anchors, to incorporate such position-matching explicitly during alignment. Here we introduce position anchoring in the alignment tool AlignMe as an aid to pairwise sequence alignment of membrane proteins. Applying this approach to realistic scenarios involving distantly-related and low complexity sequences, we illustrate how the addition of even a single anchor can dramatically improve the accuracy of the alignments, while maintaining the reproducibility and rigor of the overall alignment.

bioinformatics