bioRxiv Science⌕ Search

Biology subjects

Swairjo, M.

Publications and source records attributed to Swairjo, M..

3 recordsLinked to original sources

tRNA Modification Landscapes in Streptococci: Shared Losses and Clade-Specific Adaptations

tRNA modifications are central to bacterial translational control. Here, we integrated genetics, mass spectrometry, epitranscriptomics, and comparative genomics to map the tRNA modification genes of the Gram-positive pathogens Streptococcus mutans and Streptococcus pneumoniae. Both species show a marked loss of modifications dependent on Fe-S enzymes, consistent with a broader trend of Fe-S enzyme reduction in Streptococcus central metabolism. In addition, the D, m1A, m7G, t6A, and i6A modifications were mapped in S. pneumoniae tRNAs, and we confirmed that a unique DusB1 enzyme is responsible for the insertion of all the detectable D modifications. We uncovered differences in queuosine (Q) metabolism: while S. mutans synthesizes Q de novo, S. pneumoniae instead salvages preQ and accumulates the epoxy-Q precursor, a strategy shared with multiple other Streptococci as revealed by analysis of Q pathways in 1,599 sequenced streptococcal genomes. Comparative essentiality profiling of modification genes revealed notable differences, including the essentiality of the NLJ-threonylcarbamoyladenosine (tLJA) synthesis enzyme TsaE in S. pneumoniae but not in S. mutans, which was confirmed by genetic studies. We found that suppressor mutations in asnS encoding asparaginyl-tRNA synthetase (AsnRS) restored viability to {Delta}tsaE mutants, albeit with reduced growth. Our finding highlights the functional importance of modifications in the recognition of tRNAs by aminoacyl-tRNA synthetases.

microbiology↗

Limitations of Current Machine-Learning Models in Predicting Enzymatic Functions for Uncharacterized Proteins

Thirty to seventy percent of proteins in any given genome have no assigned function and have been labeled as the protein "unknome". This large knowledge shortfall is one of the final frontiers of biology. Machine-Learning (ML) approaches are enticing, with early successes demonstrating the ability to propagate functional knowledge from experimentally characterized proteins. An open question is the ability of machine-learning approaches to predict enzymatic functions unseen in the training sets. Using a set of Escherichia coli unknowns, we evaluated the current state-of-the-art machine-learning approaches and found that these methods currently lack the ability to integrate scientific reasoning into their prediction algorithms. While human annotators can leverage the plethora of genomic data in making plausible predictions into the unknown, current ML methods not only fail to make novel predictions but also make basic logic errors in their predictions. This underscores the need to include assessments of prediction uncertainty in model output and to test for hallucinations (logic failures) as a part of model evaluation. Explainable AI (XAI) analysis can be used to identify indicators of prediction errors, potentially identifying the most relevant data to include in the next generation of computational models. Article SummaryMany proteins in any genome, ranging from 30% to 70% of the genome, lack an assigned function. This knowledge gap limits the full use of the vast available genomic data. Machine learning has shown promise in transferring functional knowledge within isofunctional families, but it largely fails to predict novel functions not seen in its training data. Understanding these failures can guide the development of better machine-learning methods to help experts make accurate functional predictions for uncharacterized proteins.

bioinformatics↗

On the necessity to include multiple types of evidence when predicting molecular function of proteins

Machine learning-based platforms are currently revolutionizing many fields of molecular biology including structure prediction for monomers or complexes, predicting the consequences of mutations, or predicting the functions of proteins. However, these platforms use training sets based on currently available knowledge and, in essence, are not built to discover novelty. Hence, claims of discovering novel functions for protein families using artificial intelligence should be carefully dissected, as the dangers of overpredictions are real as we show in a detailed analysis of the prediction made by Kim et al 1 on the function of the YciO protein in the model organism Escherichia coli.

biochemistry↗