bioRxiv Science⌕ Search

Biology subjects

Nazia, S. Z.

Publications and source records attributed to Nazia, S. Z..

3 recordsLinked to original sources

Predicting enzyme substrate chemical structure with protein language models

The number of unannotated or orphan enzymes vastly outnumber those for which the chemical structure of the substrates are known. While a number of enzyme function prediction algorithms exist, these often predict Enzyme Commission (EC) numbers or enzyme family, which limits their ability to generate experimentally testable hypotheses. Here, we harness protein language models, cheminformatics, and machine learning classification techniques to accelerate the annotation of orphan enzymes by predicting their substrates chemical structural class. We use the orphan enzymes of Mycobacterium tuberculosis as a case study, focusing on two protein families that are highly abundant in its proteome: the short-chain dehydrogenase/reductases (SDRs) and the S-adenosylmethionine (SAM)-dependent methyltransferases. Training machine learning classification models that take as input the protein sequence embeddings obtained from a pre-trained, self-supervised protein language model results in excellent accuracy for a wide variety of prediction tasks. These include redox cofactor preference for SDRs; small-molecule vs. polymer (i.e. protein, DNA or RNA) substrate preference for SAM-dependent methyltransferases; as well as more detailed chemical structural predictions for the preferred substrates of both enzyme families. We then use these trained classifiers to generate predictions for the full set of unannotated SDRs and SAM-methyltransferases in the proteomes of M. tuberculosis and other mycobacteria, generating a set of biochemically testable hypotheses. Our approach can be extended and generalized to other enzyme families and organisms, and we envision it will help accelerate the annotation of a large number of orphan enzymes. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=106 SRC="FIGDIR/small/509940v3_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@1dab89forg.highwire.dtl.DTLVardef@8eec65org.highwire.dtl.DTLVardef@141ff13org.highwire.dtl.DTLVardef@1d16212_HPS_FORMAT_FIGEXP M_FIG C_FIG

microbiology↗

Genome-wide co-essentiality analysis in Mycobacterium tuberculosis reveals an itaconate defense enzyme module

Genome-wide random mutagenesis screens using transposon sequencing (TnSeq) have been a cornerstone of functional genetics in Mycobacterium tuberculosis (Mtb), helping to define gene essentiality across a wide range of experimental conditions. Here, we harness a recently compiled TnSeq database to identify pairwise correlations of gene essentiality profiles (i.e. co-essentiality analysis) across the Mtb genome and reveal clusters of genes with similar function. We describe selected modules identified by our pipeline, review the literature supporting their associations, and propose hypotheses about novel associations. We focus on a cluster of seven enzymes for experimental validation, characterizing it as an enzymatic arsenal that helps Mtb counter the toxic effects of itaconate, a host-derived antibacterial compound. We extend the use of these correlations to enable prediction of protein complexes by designing a virtual screen that ranks potentially interacting heterodimers from co-essential protein pairs. We envision co-essentiality analysis will help accelerate gene functional discovery in this important human pathogen.

microbiology↗

Exploring protein sequence similarity with Protein Language UMAPs (PLUMAPs)

Visualizing relationships and similarities between proteins can reveal insightful biology. Current approaches to visualize and analyze proteins based on sequence homology, such as sequence similarity networks (SSNs), create representations of BLAST-based pairwise comparisons. These approaches could benefit from incorporating recent protein language models, which generate high-dimensional vector representations of protein sequences from self-supervised learning on hundreds of millions of proteins. Inspired by SSNs, we developed an interactive tool - Protein Language UMAPs (PLUMAPs) - to visualize protein similarity with protein language models, dimensionality reduction, and topic modeling. As a case study, we compare our tool to Sequence Similarity Network (SSN) using the proteomes of two related bacterial species, Mycobacterium tuberculosis and Mycobacterium smegmatis. Both SSNs and PLUMAPs generate protein clusters corresponding to protein families and highlight enrichment or depletion across species. However, only in PLUMAPs does the layout distance between proteins and protein clusters meaningfully reflect similarity. Thus in PLUMAPs, related protein families are displayed as nearby clusters, and larger-scale structures correlate with cellular localization. Finally, we adapt techniques from topic modeling to automatically annotate protein clusters, making them more easily interpretable and potentially insightful. We envision that as large protein language models permeate bioinformatics and interactive sequence analysis tools, PLUMAPs will become a useful visualization resource across a wide variety of biological disciplines. Anticipating this, we provide a prototype for an online, open source version of PLUMAPs.

bioinformatics↗