bioRxiv Science⌕ Search

Biology subjects

Petrey, D.

Publications and source records attributed to Petrey, D..

2 recordsLinked to original sources

PrePPI: A structure informed proteome-wide database of protein-protein interactions

We present an updated version of the Predicting Protein-Protein Interactions (PrePPI) webserver which predicts PPIs on a proteome-wide scale. PrePPI combines structural and non-structural clues within a Bayesian framework to compute a likelihood ratio (LR) for essentially every possible pair of proteins in a proteome; the current database is for the human interactome. The structural modeling (SM) clue is derived from templatebased modeling and its application on a proteome-wide scale is enabled by a unique scoring function used to evaluate a putative complex. The updated version of PrePPI leverages AlphaFold structures that are parsed into individual domains. As has been demonstrated in earlier applications, PrePPI performs extremely well as measured by receiver operating characteristic curves derived from testing on E. coli and human protein-protein interaction (PPI) databases. A PrePPI database of ~1.3 million human PPIs can be queried with a webserver application that comprises multiple functionalities for examining query proteins, template complexes, 3D models for predicted complexes, and related features (https://honiglab.c2b2.columbia.edu/PrePPI). PrePPI is a state-of- the-art resource that offers an unprecedented structure-informed view of the human interactome. Graphic Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=68 SRC="FIGDIR/small/530276v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@b66c15org.highwire.dtl.DTLVardef@71d817org.highwire.dtl.DTLVardef@21d984org.highwire.dtl.DTLVardef@4f6f55_HPS_FORMAT_FIGEXP M_FIG C_FIG

systems biology↗

PrePCI: A structure- and chemical similarity-informed database of predicted protein compound interactions

We describe the Predicting Protein Compound Interactions (PrePCI) database which comprises over 5 billion predicted interactions between nearly 7 million chemical compounds and 19,797 human proteins. PrePCI relies on a proteome-wide database of structural models based on both traditional modeling techniques and the AlphaFold Protein Structure Database. Sequence and structural similarity-based metrics are established between template proteins in the Protein Data Bank, T, that bind small molecules, C, and proteins in the models database, Q. When these metrics pass a sequence threshold value, it is assumed that C also binds to Q with a probability derived from machine learning. If the relationship is based on structure, this probability is based on a scoring function that measures the extent to which C is compatible with the binding site of Q as described in the LT-scanner algorithm. For every predicted complex derived in this way, chemical similarity based on the Tanimoto Coefficient identifies other small molecules that may bind to Q. A likelihood ratio for the binding of C to Q is obtained from naive Bayesian statistics. The PrePCI algorithm performs well under different validations. It can be queried by entering a UniProt ID for a protein and obtaining a list of compounds predicted to bind to it along with associated probabilities. Alternatively, entering an identifier for the compound outputs a list of proteins it is predicted to bind. Specific applications of the database are described and a strategy is introduced to use PrePCI as a first step in a docking screen.

systems biology↗