bioRxiv Science⌕ Search

Biology subjects

Waman, V.

Publications and source records attributed to Waman, V..

2 recordsLinked to original sources

Understanding structural and functional diversity of ATP-PPases using protein domains and functional families in CATH database

ATP-Pyrophosphatases (ATP-PPases) are the most primordial lineage of the large and diverse HUP (HIGH-motif proteins, Universal Stress Proteins, ATP-Pyrophosphatase) superfamily. There are four different ATP-PPase substrate-specificity groups, and members of each group show considerable sequence variation across the domains of life despite sharing the same catalytic function. Over the past decade, there has been a >20-fold expansion in the number of ATP-PPase domain structures most recently from advances in protein structure prediction (e.g. Alphafold2). Using the enriched structural information, we have characterised the two most populated ATP-PPase substrate-specificity groups, the NAD-synthases (NAD) and GMP synthases (GMPS). We performed local structural and sequence comparisons between the NADS and GMPS from different domains of life and identified taxonomic-group specific structural functional motifs. As GMPS and NADS are potential drug targets of pathogenic microorganisms including Mycobacterium tuberculosis, structural motifs specific to bacterial GMPS and NADS provide new insights that may aid antibacterial-drug design.

bioinformatics↗

CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models

1.CATH is a protein domain classification resource that combines an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues that might be missed by state-of-the-art HMM-based approaches. The proposed algorithm for this task (CATHe) combines a neural network with sequence representations obtained from protein language models. The employed dataset consisted of remote homologues that had less than 20% sequence identity. The CATHe models trained on 1773 largest, and 50 largest CATH superfamilies had an accuracy of 85.6+-0.4, and 98.15+-0.30 respectively. To examine whether CATHe was able to detect more remote homologues than HMM-based approaches, we employed a dataset consisting of protein regions that had annotations in Pfam, but not in CATH. For this experiment, we used highly reliable CATHe predictions (expected error rate <0.5%), which provided CATH annotations for 4.62 million Pfam domains. For a subset of these domains from homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold structures with experimental structures from the CATHe predicted superfamilies.

bioinformatics↗