bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.11.19.689359

A Practical Resource for Multi-Omics Data Integration in Microbial Systems

Abstract

The increasing availability of microbial multi-omics datasets has created new opportunities to explore complex biological systems. However, exploration remains limited by the lack of accessible, reproducible workflows that integrate multiple omics layers and deliver easily interpretable visualisations of functional and pathway-level insights. Here, we present an R-based workflow for integrated analysis and network-based pathway visualisation of microbial multi-omic data. The workflow enables microbiologists to analyse transcriptomic, proteomic, and metabolomic datasets either individually or in combination, apply univariate and multivariate approaches for biomarker discovery, and generate easily interpretable visualisations of functional and pathway-level signatures. Implemented as multi-step R Markdown, it leverages widely-used open-source tools, including mixOmics for biomarker identification and omics integration and clusterProfiler for pathway and functional enrichment analyses, with a new network-based integration and visualisation. Its flexible design supports a range of experimental structures and facilitates comparisons across strains, omics layers, and conditions, making it suitable for researchers with limited computational expertise. We demonstrate its utility using a publicly available Streptococcus pyogenes dataset, revealing both shared and strain-specific functional responses to human serum. This workflow provides a comprehensive and adaptable framework for systematic multi-omics analysis, improving accessibility and reproducibility and facilitating functional interpretation of microbial responses to diverse environments. Data summaryThe code for this workflow is available on GitHub (https://github.com/warasinee/Multiomics_Case_Study). Datasets from our previously published study (1) were used to showcase the functionality and practical utility of the workflow. The multi-omics Streptococcus pyogenes dataset used in this study is available in the following public repositories: Gene Expression Omnibus (GSE152821; GSE152822; GSE152823; GSE152824, GSE152826), Proteomics Identifications Database (PXD020863), and MetaobLights (MTBLS2324) (1). Impact StatementHigh-throughput omics technologies are transforming our understanding of how microbes adapt to diverse environments and cause disease. The integration of diverse omics layers at a systems level, combining transcriptomics, proteomics, and metabolomics data to identify signature molecules, pathways, and their interactions, remains challenging. Here, we present an R-based bioinformatic workflow designed for microbiology research, which connects existing tools and customised functions to streamline data integration and interpretation. The workflow links biomarkers to functional pathways, visualises results in an interactive network context, and is designed for flexibility and reproducibility. This practical resource lowers technical barriers to microbial multi-omics analysis, providing user-friendly access for dataset exploration and integration and supporting interpretation of system-level microbial adaptation in environmental, clinical, and industrial contexts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mujchariyakul, W., Hachani, A., Stinear, T. P., Le Cao, K.-A., Howden, B. P., Walsh, C. J., Guerillot, R.. 2025-11-19. A Practical Resource for Multi-Omics Data Integration in Microbial Systems. https://doi.org/10.1101/2025.11.19.689359

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Limit-pushing overexpression reveals constraints on protein abundance

Proteins are often classified as toxic or non-toxic without measuring the abundance reached, leaving constraints on tolerable protein abundance unresolved. We established a limit-pushing approach in Saccharomyces cerevisiae combining strong inducible expression with gTOW-mediated high-copy selection to counteract copy-number compensation while measuring protein abundance and growth. Nearly all of approximately 80 chromosome I proteins severely inhibited growth or reduced viability at sufficiently high abundance. We established IE50, the expression level associated with a 50% reduction in growth rate, to quantify their widely varying overexpression tolerance. IE50 was positively associated with predicted structural order and cytoplasmic localization propensity and negatively associated with sulphur content. Single-cell imaging linked higher tolerance to proteins remaining cytoplasmic without becoming aggregation-positive and revealed abundance-dependent changes in localization and organelle morphology. At extreme abundance, Fun12, Nup60, and Pex22 generated distinct large-scale intracellular states through specific sequence regions. These findings establish overexpression toxicity as a quantitative property linked to protein characteristics and reveal both constraints on tolerable abundance and sequence-dependent capacities for intracellular organization.

systems biology↗

Accessing Enzyme Kinetic Data and Prediction Methods at Scale

Enzyme kinetic parameters inform metabolic models, yet experimental measurements are sparse. A growing body of work predicts them from protein and substrate features, but software fragmentation hinders adoption, so downstream tools lock into the most accessible method. We present OpenKinetics Predictor (at predictor.openkinetics.org), an open-source platform integrating thirteen methods in isolated environments behind one interface. The platform optionally reports similarity between query proteins and each method's training data to contextualise reliability. A common featurisation-prediction abstraction keeps it extensible, and independent parties, including original authors, contributed many methods. We pair it with a data portal (at data.openkinetics.org) that exposes CatLog, a curated kinetic dataset, with precomputed embeddings, predicted binding sites, and standardised splits. Both offer a web interface and an API, and the GECKO modelling toolbox already calls the predictor API. As a case study, we predict across an E. coli model and find inter-predictor agreement varies with metabolic context and data availability.

systems biology↗

A thermoregulatory design principle for transitions into hypometabolism

Mammals entering torpor or hibernation undergo an abrupt transition from normothermia to hypothermia, yet how thermoregulation enables this switch remains poorly understood. Here, we identify dynamical signatures that precede these transitions and a mathematical principle that can generate them. In fasting-induced torpor in mice, body-temperature fluctuations increased before torpor onset, providing an early-warning signal that tracked proximity to the transition better than temperature decline alone. A heat-balance model showed that reducing how strongly the effective heat-loss coefficient depends on body temperature reorganizes thermoregulatory stability, allowing a low-temperature equilibrium to emerge while the normothermic state remains stable. This organization is consistent with a symmetry-broken pitchfork involving a saddle-node. Similar increases in temperature fluctuations preceded hibernation onset in hamsters. These findings link pre-transition temperature dynamics to changes in the underlying thermoregulatory landscape and provide a framework for detecting and understanding transitions from normothermia to hypothermia.

systems biology↗