bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.07.15.203653

Annotation of ribosomal protein mass peaks in MALDI-TOF mass spectra of bacterial species and their phylogenetic significance

Abstract

Although MALDI-TOF mass spectrometry based microbial identification has achieved a level of accuracy that facilitate its use in classifying microbes to the species and strain level, questions remain on the identities of the mass peaks profiled from individual microbial species. Specifically, in the popular approach of comparing the mass spectrum of known and unknown microbes for identification purposes, the identities of the mass peaks were not taken into consideration. This study sought to determine if ribosomal proteins could account for some of the mass peaks profiled in MALDI-TOF mass spectra of different bacterial species. Using calculated molecular mass of ribosomal proteins for annotating mass peaks in bacterial species MALDI-TOF mass spectra downloaded from the SpectraBank database, this study revealed that ribosomal proteins could account for the low molecular weight mass peaks of <10000 Da. However, contrary to published reports, ribosomal proteins could not account for most of the mass peaks profiled. In particular, the data revealed that between 1 and 6 ribosomal protein mass peaks could be annotated in each mass spectrum. Annotated ribosomal proteins were S16, S17, S18, S20 and S21 from the small ribosome subunit, and L27, L28, L29, L30, L31, L31 Type B, L32, L33, L34, L35 and L36 from the large ribosome subunit. The ribosomal proteins with the most number of mass peak annotations were L36 and L29, with L34, L33, and L31 completing the list of ribosomal proteins with large number of annotations. Given the highly conserved nature of most ribosomal proteins, possible phylogenetic significance of the annotated ribosomal proteins were investigated through reconstruction of maximum likelihood phylogenetic trees. Results revealed that except for ribosomal protein L34, L31, L36 and S18, all annotated ribosomal proteins hold phylogenetic significance under the criteria of recapitulation of phylogenetic cluster groups present in the phylogeny of 16S rRNA. Phylogenetic significance of the annotated ribosomal proteins was further verified by the phylogenetic tree constructed based on the concatenated amino acid sequence of L29, S16, S20, S17, L27 and L35. Finally, analysis of the structure of the annotated ribosomal proteins did not reveal a high conservation of structure of the ribosomal proteins. Collectively, small low molecular weight (<10000 Da) ribosomal proteins could annotate some of the mass peaks in MALDI-TOF mass spectra of various bacterial species, and most of the ribosomal proteins hold phylogenetic significance. However, structural analysis did not identify a conserved structure for the annotated ribosomal proteins. Annotation of ribosomal protein mass peaks in MALDI-TOF mass spectra highlighted the deep biological basis inherent in the mass spectrometry-based microbial identification method. Subject areas biochemistry, biotechnology, microbiology, evolution, ecology Significance of the workWhile MALDI-TOF MS has been successfully used in identification of different microbes to the species and strain level through the comparison of mass spectra of known and unknown microbes, the approach (known as mass spectrum fingerprinting) remains lacking in the biological basis that underpins the technique. This study sought to uncover some of the biological basis that underpins MALDI-TOF MS microbial identification through the annotation of profiled mass peaks with ribosomal proteins. Previous studies have linked different ribosomal proteins to mass peaks in MALDI-TOF mass spectra of bacteria; however, broad spectrum verification of the finding across multiple species across different genera remain lacking. Using a collection of MALDI-TOF mass spectra of 110 bacterial species and strains catalogued in SpectraBank, this study sought to annotate ribosomal protein mass peaks in the mass spectra. Results revealed that small, low molecular weight ribosomal proteins of molecular mass < 10000 Da could annotate between 1 and 6 mass peaks in the catalogued mass spectra. This was smaller than the number of ribosomal proteins mass peaks postulated by previous studies. Overall, 16 ribosomal proteins (S16, S17, S18, S20, S21, L27, L28, L29, L30, L31, L31 Type B, L32, L33, L34, L35, and L36) were annotated with the most number of mass peaks annotations coming from L36 and L29. Reconstruction of phylogenetic trees of the annotated ribosomal proteins revealed that most of the ribosomal proteins hold phylogenetic significance with respect to the phylogeny of 16S rRNA. This provided further evidence that a deep biological basis is present in the approach of using mass spectrometry profiling of biomolecules for identifying bacterial species. HighlightsO_LIRibosomal protein mass peaks were annotated in MALDI-TOF mass spectra of bacterial species across multiple genera. C_LIO_LIAnnotated ribosomal proteins were S16, S17, S18, S20, S21 for the small ribosome subunit, and L27, L28, L29, L30, L31, L31 Type B, L32, L33, L34, L35, L36 for the large ribosome subunit. C_LIO_LIBetween 1 and 6 ribosomal protein mass peaks were annotated per mass spectrum, a number significantly lower than that implied by other studies. C_LIO_LIAnnotated ribosomal proteins were small, low molecular weight ribosomal proteins of molecular mass < 10000 Da. C_LIO_LIPhylogenetic tree reconstruction revealed the phylogenetic significance of most annotated ribosomal proteins except ribosomal protein L34, L31, L36 and S18. C_LIO_LIMulti-locus sequence typing of L29, S16, S20, S17, L27 and L35 further showed the phylogenetic significance of ribosomal proteins in recapitulating the phylogeny of 16S rRNA. C_LIO_LIStructural analysis of annotated ribosomal proteins did not find conserved structure. Thus, the reasons for the annotation of particular ribosomal proteins over others remain unknown. C_LI

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ng, W.. 2020-07-15. Annotation of ribosomal protein mass peaks in MALDI-TOF mass spectra of bacterial species and their phylogenetic significance. https://doi.org/10.1101/2020.07.15.203653

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

spatialMET: an open and scalable framework for spatial metabolomics analysis

Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.

bioinformatics

Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach

Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.

bioinformatics

An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.

Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.

bioinformatics