bioRxiv Science⌕ Search

Biology subjects

Faezov, B.

Publications and source records attributed to Faezov, B..

6 recordsLinked to original sources

AlphaFold2 models of the active form of all 437 catalytically-competent typical human kinase domains

AbstractHumans have 437 catalytically competent protein kinase domains with the typical kinase fold, similar to the structure of Protein Kinase A (PKA). Only 155 of these kinases are in the Protein Data Bank in their active form. The active form of a kinase must satisfy requirements for binding ATP, magnesium, and substrate. From structural bioinformatics analysis of 40 unique substrate-bound kinases, we derived several criteria for the active form of protein kinases. We include requirements on the DFG motif of the activation loop but also on the positions of the N-terminal and C-terminal segments of the activation loop that must be placed appropriately to bind substrate. Because the active form of catalytic kinases is needed for understanding substrate specificity and the effects of mutations on catalytic activity in cancer and other diseases, we used AlphaFold2 to produce models of all 437 human protein kinases in the active form. This was accomplished with templates in the active form from the PDB and shallow multiple sequence alignments of orthologs and close homologs of the query protein. We selected models for each kinase based on the pLDDT scores of the activation loop residues, demonstrating that the highest scoring models have the lowest or close to the lowest RMSD to 22 non-redundant substrate-bound structures in the PDB. A larger benchmark of all 130 active kinase structures with complete activation loops in the PDB shows that 80% of the highest-scoring AlphaFold2 models have RMSD < 1.0 [A] and 90% have RMSD < 2.0 [A] over the activation loop backbone atoms. Models for all 437 catalytic kinases are available at http://dunbrack.fccc.edu/kincore/activemodels. We believe they may be useful for interpreting mutations leading to constitutive catalytic activity in cancer as well as for templates for modeling substrate and inhibitor binding for molecules which bind to the active state.

bioinformatics↗

FAM122A ensures cell cycle interphase progression and checkpoint control as a SLiM-dependent substrate-competitive inhibitor to the B55/PP2A phosphatase

The Ser/Thr protein phosphatase 2A (PP2A) is a highly conserved collection of heterotrimeric holoenzymes responsible for the dephosphorylation of many regulated phosphoproteins. Substrate recognition and the integration of regulatory cues are mediated by B regulatory subunits that are complexed to the catalytic subunit (C) by a scaffold protein (A). PP2A/B55 substrate recruitment was thought to be mediated by charge-charge interactions between the surface of B55 and its substrates. Challenging this view, we recently discovered a conserved SLiM [RK]-V-x-x-[VI]-R in a range of proteins, including substrates such as the retinoblastoma-related protein p107 and TAU (Fowle et al. eLife 2021;10:e63181). Here we report the identification of this SLiM in FAM122A, an inhibitor of B55/PP2A. This conserved SLiM is necessary for FAM122A binding to B55 in vitro and in cells. Computational structure prediction with AlphaFold2 predicts an interaction consistent with the mutational and biochemical data and supports a mechanism whereby FAM122A uses the SLiM in the form of a short -helix to dock to the B55 top groove. In this model, FAM122A spatially constrains substrate access by occluding the catalytic subunit with a second -helix immediately adjacent to helix 1. Consistently, FAM122A functions as a competitive inhibitor as it prevents binding of substrates in in vitro competition assays and the dephosphorylation of CDK substrates by B55/PP2A in cell lysates. Ablation of FAM122A in human cell lines reduces the rate of proliferation, progression through cell cycle transitions and abrogates G1/S and intra-S phase cell cycle checkpoints. FAM122A-KO in HEK293 cells results in attenuation of CHK1 and CHK2 activation in response to replication stress. Overall, these data strongly suggest that FAM122A is a SLiM-dependent, substrate-competitive inhibitor of B55/PP2A that suppresses multiple functions of B55 in the DNA damage response and in timely progression through the cell cycle interphase.

molecular biology↗

Co-occurring mutations in the POLE exonuclease and non-exonuclease domains define a unique subset of highly mutagenic tumors.

Somatic POLE mutations in the exonuclease domain (ExoD) are prevalent in colorectal cancer (CRC), endometrial cancer (EC), and others and typically lead to dramatically increased tumor mutation burden (TMB). To understand whether non-ExoD mutations also play a role in mutagenesis, we assessed TMB in 447/14541 POLE-mutated CRCs, ECs, and ovarian cancers (OC) based on classification TMB-High (TMB-H) or TMB-Low (TMB-L). TMB-H tumors were segregated as POLE ExoD driver, POLE ExoD driver plus POLE Variant, and POLE Variant TMB-H. Intriguingly, TMB was highest in tumors bearing POLE ExoD driver plus POLE Variant (p<0.001 in CRC and EC, p<0.05 in OC). Integrated analysis of AlphaFold2-modeled POLE models and quantitative estimate of stability indicated that multiple variants had significant impact on functionality. These data indicate that co-occurring POLE variants categorize a unique subset of POLE-driven tumors defined by ultra-high TMB, which has implications for abundance of tumor neoantigens, therapeutic response, and patient outcomes. SignificanceSomatic POLE ExoD driver mutations cause proofreading deficiency that induces high tumor mutation burden (TMB). This study defines a novel modifier role for non-ExoD mutations in POLE ExoD-driven tumors, associated with ultra-high TMB. These data may inform acquisition of tumor neoantigens, tumor classification, therapeutic response, and patient outcomes.

cancer biology↗

A penultimate classification of canonical antibody CDR conformations

Antibody complementarity determining regions (CDRs) are loops within antibodies responsible for engaging antigens during the immune response and in antibody therapeutics and laboratory reagents. Since the 1980s, the conformations of the hypervariable CDRs have been structurally classified into a number of "canonical conformations" by Chothia, Lesk, Thornton, and others. In 2011 (North et al, J Mol Biol. 2011), we produced a quantitative clustering of approximately 300 structures of each CDR based on their length, a dihedral angle metric, and an affinity propagation algorithm. The data have been made available on our PyIgClassify website since 2015 and have been widely used in assigning conformational labels to antibodies in new structures and in molecular dynamics simulations. In the years since, it is has become apparent that many of the clusters are not "canonical" since they have not grown in size and still contain few sequences. Some clusters represent multiple conformations, given the assignment method we have used since 2015. Electron density calculations indicate that some clusters are due to misfitting of coordinates to electron density. In this work, we have performed a new statistical clustering of antibody CDR conformations. We used Electron Density in Atoms (EDIA, Meyder et al., 2017) to produce data sets with different levels of electron density validation. Clusters were chosen by their presence in high electron density cutoff data sets and with sufficient sequences ([&ge;]10) across the entire PDB (no EDIA cutoff). About half of the North et al. clusters have been "retired" and 13 new clusters have been identified. We also include clustering of the H4 and L4 CDRs, otherwise known as the "DE loop" which connects strands D and E of the variable domain. The DE loop sometimes contacts antigens and affects the structure of neighboring CDR1 and CDR2 loops. The current database contains 6,486 PDB antibody entries. The new clustering will be useful in the analysis and development of new antibody structure prediction and design algorithms based on rapidly emerging techniques in deep learning. The new clustering data are available at http://dunbrack2.fccc.edu/PyIgClassify2.

immunology↗

Development and utility of a PAK1-selective degrader

Amplification and/or overexpression of the PAK1 gene is common in several malignancies, and inhibition of PAK1 by small molecules has been shown to impede the growth and survival of such cells. Potent inhibitors of PAK1 and its close relatives, PAK2, and PAK3, have been described, but clinical development has been hindered by recent findings that PAK2 function is required for normal cardiovascular function in adult mice. A unique allosteric PAK1-selective inhibitor, NVS-PAK1-1, provides a potential path forward, but has relatively modest potency in cells. Here, we report the development of BJG-05-039, a PAK1-seletive degrader consisting of the allosteric PAK1 inhibitor NVS-PAK1-1 conjugated to lenalidomide, a recruiter of the E3 ubiquitin ligase substrate adaptor Cereblon (CRBN). BJG-05-039 induced degradation of PAK1, but not PAK2, and displayed enhanced anti-proliferative effects relative to its parent compound in PAK1-dependent, but not PAK2-dependent, cell lines. Notably, BJG-05-039 promoted sustained PAK1 degradation and inhibition of downstream signaling effects at ten-fold lower dosage than NVS-PAK1-1. Our findings suggest that selective PAK1 degradation may confer more potent pharmacological effects compared with catalytic inhibition and highlight the potential advantages of PAK1-targeted degradation.

cancer biology↗

PDBrenum: a webserver and program providing Protein Data Bank files renumbered according to their UniProt sequences

The Protein Data Bank (PDB) was established at Brookhaven National Laboratories in 1971 as an archive for biological macromolecular crystal structures. In early 2021, the database has more than 175,000 structures solved by X-ray crystallography, nuclear magnetic resonance, cryo-electron microscopy, and other methods. Many proteins have been studied under different conditions, including binding partners such as ligands, nucleic acids, or other proteins; mutations, and post-translational modifications, thus enabling extensive comparative structure-function studies. However, these studies are made more difficult because authors are allowed by the PDB to number the amino acids in each protein sequence in any manner they wish. This results in the same protein being numbered differently in the available PDB entries. For instance, some authors may include N-terminal signal peptides or the N-terminal methionine in the sequence numbering and others may not. In addition to the coordinates, there are many fields that contain information regarding specific residues in the sequence of each protein in the entry. Here we provide a webserver and Python3 application that fixes the PDB sequence numbering problem by replacing the author numbering with numbering derived from the corresponding UniProt sequences. We obtain this correspondence from the SIFTS database from PDBe. The server and program can take a list of PDB entries or a list of UniProt identifiers (e.g., "P04637" or "P53_HUMAN") and provide renumbered files in mmCIF format and the legacy PDB format for both asymmetric unit files and biological assembly files provided by PDBe. AvailabilitySource code is freely available at https://github.com/Faezov/PDBrenum. The webserver is located at: http://dunbrack3.fccc.edu/PDBrenum. Contactbulat.faezov@fccc.edu or roland.dunbrack@fccc.edu.

bioinformatics↗