bioRxiv Science⌕ Search

Biology subjects

Tellez, A. V.

Publications and source records attributed to Tellez, A. V..

2 recordsLinked to original sources

Predicting enzyme substrate chemical structure with protein language models

The number of unannotated or orphan enzymes vastly outnumber those for which the chemical structure of the substrates are known. While a number of enzyme function prediction algorithms exist, these often predict Enzyme Commission (EC) numbers or enzyme family, which limits their ability to generate experimentally testable hypotheses. Here, we harness protein language models, cheminformatics, and machine learning classification techniques to accelerate the annotation of orphan enzymes by predicting their substrates chemical structural class. We use the orphan enzymes of Mycobacterium tuberculosis as a case study, focusing on two protein families that are highly abundant in its proteome: the short-chain dehydrogenase/reductases (SDRs) and the S-adenosylmethionine (SAM)-dependent methyltransferases. Training machine learning classification models that take as input the protein sequence embeddings obtained from a pre-trained, self-supervised protein language model results in excellent accuracy for a wide variety of prediction tasks. These include redox cofactor preference for SDRs; small-molecule vs. polymer (i.e. protein, DNA or RNA) substrate preference for SAM-dependent methyltransferases; as well as more detailed chemical structural predictions for the preferred substrates of both enzyme families. We then use these trained classifiers to generate predictions for the full set of unannotated SDRs and SAM-methyltransferases in the proteomes of M. tuberculosis and other mycobacteria, generating a set of biochemically testable hypotheses. Our approach can be extended and generalized to other enzyme families and organisms, and we envision it will help accelerate the annotation of a large number of orphan enzymes. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=106 SRC="FIGDIR/small/509940v3_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@1dab89forg.highwire.dtl.DTLVardef@8eec65org.highwire.dtl.DTLVardef@141ff13org.highwire.dtl.DTLVardef@1d16212_HPS_FORMAT_FIGEXP M_FIG C_FIG

microbiology↗

Genome-wide co-essentiality analysis in Mycobacterium tuberculosis reveals an itaconate defense enzyme module

Genome-wide random mutagenesis screens using transposon sequencing (TnSeq) have been a cornerstone of functional genetics in Mycobacterium tuberculosis (Mtb), helping to define gene essentiality across a wide range of experimental conditions. Here, we harness a recently compiled TnSeq database to identify pairwise correlations of gene essentiality profiles (i.e. co-essentiality analysis) across the Mtb genome and reveal clusters of genes with similar function. We describe selected modules identified by our pipeline, review the literature supporting their associations, and propose hypotheses about novel associations. We focus on a cluster of seven enzymes for experimental validation, characterizing it as an enzymatic arsenal that helps Mtb counter the toxic effects of itaconate, a host-derived antibacterial compound. We extend the use of these correlations to enable prediction of protein complexes by designing a virtual screen that ranks potentially interacting heterodimers from co-essential protein pairs. We envision co-essentiality analysis will help accelerate gene functional discovery in this important human pathogen.

microbiology↗