bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.02.23.707520

What microbes want: exploring microbial substrate preferences with the Web of Microbes Agent

Abstract

Understanding and predicting bacterial substrate preferences has broad utility from microbial interactions to selecting prebiotics. Isolate exometabolite profiling directly measures which compounds a given microbe utilizes from an array of metabolites in the environment. However, modeling, mining, and integrating these data are challenging. Here, we introduce a Bayesian Personalized Ranking (BPR) model applied to substrate preferences which we find learns to rank compounds by a given microbes preference. It was found to outperform the other ranking models (AUC = 0.93), proved robust to ablation, showed strong within-genus isolate pairs correlation (Spearman rank = 0.78) and predictive ability for new data. BPR was then used to create the Web of Microbes (WoM) Agent by integrating it with the Phydon growth model and Large Language Model (LLM) for autonomous orchestration tool calling and analysis. The WoM Agent accurately predicted substrate consumption by existing strain grown on a novel medium and correctly identified bacteria enriched in soil metabolite spike-in experiments. Additionally, the WoM Agent can use autonomous reasoning including to predict substrates that will selectively promote the growth of one clade of bacteria over another including helping interpret results and suggest new hypotheses and experiments. We anticipate broad applications in microbial cultivation, microbiome engineering, and environmental microbiology, with the agents capabilities further extensible through the integration of additional tools and use of rapidly improving LLMs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Northen, T. R., de Raad, M., Kosina, S. M., Andeer, P. F., Novak, V., Biggs, B., Peng, H., Paulitz, T., Arkin, A. P., Louie, K. B., Wang, M., Bowen, B. P.. 2026-02-24. What microbes want: exploring microbial substrate preferences with the Web of Microbes Agent. https://doi.org/10.64898/2026.02.23.707520

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

SpaReg: sparsity-based 3D reconstruction of tissue microenvironments at native resolution across morphological and spatial molecular modalities

Tissue microenvironments comprise cellular and acellular components whose three-dimensional (3D) architecture guides disease fate. Direct imaging of intact specimens by light-sheet and multiphoton microscopy, and computational reconstruction from serial sections, have established that 3D spatial context reveals cell and tissue organization inaccessible at single planes. Computational reconstruction in particular can leverage archived human tissue, benefiting from the cost-effectiveness, robustness, scalable storage, workflow compatibility, and century-long pathobiology knowledge of histology, and can integrate multiple spatial modalities. However, sectioning can introduce tears and folds, and computational alignment can further distort tissue integrity. Here we introduce SpaReg, a sparsity-based 3D reconstruction method spanning histology, spatial proteomics and spatial transcriptomics. Across multiple organs, SpaReg robustly reconstructs large tissue volumes with preserved subcellular morphology despite sectioning artifacts. On a standardized histology benchmark, SpaReg achieves the best balance between 3D reconstruction accuracy and tissue integrity, and on spatial transcriptomics benchmarks it ranks among the leading methods while scaling to hundreds of sections and millions of cells in a dataset that several existing methods fail to process. Preservation of subcellular morphology by SpaReg also enables training of a Hematoxylin and Eosin (H&E)-based epithelial, T and B cell classifier, generating single-cell-resolved 3D maps directly from H&E. Applied to pancreatic tissue containing pancreatic ductal adenocarcinoma arising from an intraductal papillary mucinous neoplasm, these maps reveal that 2D sections overestimate immune exclusion, and resolve lymphoid aggregates in 3D. SpaReg, therefore, provides a scalable foundation for morphologically faithful, multimodal 3D atlases and spatially informed disease modeling

systems biology↗

TxCyto: A machine learning framework for estimating cytokine activity from whole transcriptome

Cytokines are critical mediators of intercellular communication, and a comprehensive characterization of their activity is essential for understanding health and disease. Existing tools to infer cytokine activity rely on experimental measurements. However, such measurements are available only for a small minority (43) of cytokines, and moreover, cytokine activity and response are highly context-specific, making a comprehensive experimental profiling across tissues, disease states, and biological contexts impractical. To address this gap, we developed TxCyto - a deep learning-based framework that infers the activity of cytokines, and more broadly of the tumor secretome, directly from the whole transcriptome profile of a sample. Trained on pan-cancer TCGA tumor transcriptomes, TxCyto was extensively validated in multiple independent datasets, including cytokine perturbation experiments. Across multiple cancer immunotherapy cohorts, TxCyto identified cytokines whose predicted activity was associated with therapeutic response. Furthermore, in spatial transcriptomic data for Liver cancer, TxCyto discovered spatial niches associated with response to immunotherapy. Overall, we develop a machine learning tool -TxCyto, for predicting the activity of 645 cytokines and tumor secretome from readily available whole transcriptomes. The TxCyto framework is generally applicable to other classes of regulatory molecules and TxCyto code base, and the tools are provided at https://github.com/Rahulncbs/TxCyto.

systems biology↗

Interpretable machine learning coupled to gene regulatory networks uncovers subcircuits underlying cell fate decisions

Gene regulatory networks (GRNs) model causal linkages that control cell fate decisions and differentiation transitions. Prioritizing regulatory subnetworks underlying cell state differences is of critical importance, but current methods including those reliant on topological metrics introduce circularity as the metrics prioritizing TFs are computed from the same networks whose assumptions they inherit. Separately, interpretable machine learning methods can identify latent factors (LFs) that discriminate cellular states with formal statistical guarantees but do not model regulatory linkages. Here, we present FOCAL (Factor-Outcome Coupling for Assessment of Linkages), a paradigm to prioritize regulatory subnetworks by coupling state-specific and dynamic GRNs with outcome-supervised LFs learned using interpretable machine learning without reference to network topology. This shifts GRN focus from macroscopic TF nodes to state-specific and dynamic TF-gene linkages. In B and T cells, FOCAL identified GIFs (GRNs coupled to Interpretable latent Factors), prioritized regulatory subnetworks underlying established states as well as transient regulatory episodes preceding them. By coupling LFs learnt from perturbation experiments of lineage-defining TFs, FOCAL identified transcriptional predisposition to alternative fates within progenitor cell populations before overt differentiation. This uncovered a novel NFATC2-IRF8 interplay in activated B cells, that was validated by in-vitro and in-vivo genetic perturbations. The two transcription factors act cooperatively to restrain extrafollicular plasmablast differentiation and promote germinal center B cell fate.

systems biology↗