bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.05.19.726160

Benchmarking strain-level profiling of Escherichia coli in short-read gut metagenomes

Abstract

2.Metagenomes offer the potential to characterise Escherichia coli strain-level diversity within the human gut microbiome, informing our understanding of colonisation diversity and the genetic features distinguishing infection from carriage. Among numerous reference-based tools for short-read metagenomic strain-level profiling, the best approach remains unclear. Here, we benchmarked six published tools--PanTax, PathoScope, StrainGE, Strainify, StrainR2 and StrainScan--for their ability to detect co-existing strains of E. coli and estimate their relative abundance across real and simulated metagenomes of increasing complexity with varying reference database composition. In the ZymoBIOMICS(R) D6331 dataset, only PanTax achieved zero error when predicting the equal abundance of five E. coli strains. In a differentially abundant four-strain mock community dataset (SRR13355226), StrainScan had the lowest mean absolute proportional error (0.89), driven by reduced sensitivity (0.5), followed by PathoScope (4.08). Across simulated metagenomes reflecting the healthy adult gut microbiome, all tools demonstrated high sensitivity ([≥]0.833), but specificity, precision and F1 score were selectively improved in some tools through detection thresholds to remove low abundance false positives. Outright, StrainGE achieved the highest F1 score (0.978). Predicted relative abundances of the E. coli K12-MG1655 (phylogroup A) and O157:H7 Sakai (phylogroup E) strains spiked into simulated metagenomes across varying abundance ratios were generally accurate, with PanTax and StrainR2 showing the lowest mean absolute proportional error (0.06). When truly present strains were removed from the reference database, out-of-phylogroup assignments were observed for some tools. Collectively, our results demonstrate that published metagenomic strain-level profiling tools vary in their ability to profile E. coli strains, indicating that method selection should be guided by intended application. These findings will facilitate characterisation of E. coli strain-level diversity within short-read gut metagenomes with greater accuracy than previously possible. 3. Impact statementStrain-level diversity within the human gut microbiome can be important for human health, with species such as Escherichia coli existing as both commensal and pathogenic strains. Most existing gut microbiome datasets are from short-read i.e., Illumina, sequencing, and numerous bioinformatic tools have been developed to profile strain-level variation from these data. However, the existing literature is often difficult to navigate given that the available tools have been benchmarked in various ways and are subject to author bias. This is, to our knowledge, the first independent benchmarking of six published tools for profiling E. coli at strain-level resolution from short-read metagenomes. Using both real and simulated datasets of increasing complexity, we demonstrate substantial variation in tool performance in terms of strain detection and relative abundance estimation, highlighting that tool choice should be guided by the specific research question, as no single method performs optimally across all scenarios. This work provides an unbiased framework for tool selection and will support more accurate and reproducible E. coli strain-level analyses in gut microbiome research from short-metagenomic data. 4. Data summaryThe authors confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. Supplementary methods, six supplementary tables and four supplementary figures are available in the online Supplementary Material. Code for simulating metagenomes using InSilicoSeq, SLURM job scripts for the simulated metagenomes dataset and R visualization and statistical analysis scripts are available within a dedicated public GitHub repository (https://github.com/mattgal11/benchmarking_short_read_strain_profilers). The following supplementary data are available on FigShare (https://doi.org/10.6084/m9.figshare.32125474): O_LINormalised per-contig relative abundances for 98 species assemblies used to construct the baseline gut microbiome profile for InSilicoSeq metagenome simulation (Normalised_relative_abundance_for_InSilicoSeq_simulated_metagenomes_ gut_microbiome_profile.csv) C_LIO_LIZymoBIOMICS(R) D6331 gut microbiome standard dataset predicted relative abundance data (Zymobiomics_D6331_raw_predicted_abundance.csv) C_LIO_LISRR13355226 mock community (99% human reads; 1% E. coli reads) paired-end reads with human reads depleted (SRR13355226_depleted_R1.fastq.gz & SRR13355226_depleted_R2.fastq.gz) C_LIO_LISRR13355226 mock community dataset raw predicted abundance data, with and without human read removal (SRR13355226_raw_predicted_abundance_with_and_without_human_read_r emoval.csv) C_LIO_LISimulated metagenomes dataset raw call types and detection metric values with increasing detection thresholds (Simulated_metagenomes_raw_call_type_assingments_and_detection_thres holds.csv) C_LIO_LISimulated metagenomes dataset (all references) predicted relative abundance data (Simulated_metagenomes_all_references_raw_predicted_abundances.csv) C_LIO_LISimulated metagenomes dataset (all references) mapped reads for PathoScope and Strainify (all_refs_pathoscope_reads_mapped.csv & all_refs_strainify_reads_mapped.csv) C_LIO_LISimulated metagenomes dataset (reduced reference database) predicted relative abundance data (Simulated_metagenomes_K12_and_Sakai_removed_from_reference_datab ase_raw_predicted_abundance.csv) C_LI

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Galbraith, M., Williams, D., Shaw, L. P., Lipworth, S., Stoesser, N.. 2026-05-19. Benchmarking strain-level profiling of Escherichia coli in short-read gut metagenomes. https://doi.org/10.64898/2026.05.19.726160

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

BiomiX 3.0: A user-friendly platform for democratized multi-omics integration with graph-based learning.

Background Multi-omics integration has emerged as a powerful strategy to decode the molecular complexity of biological systems. However, the diversity of available methods, each designed with distinct assumptions, objectives, and computational requirements, makes method selection, usage and interpretation challenging for nonexpert users. Here we present BiomiX 3.0, an updated version of the BiomiX platform that extends its integration capabilities with four additional methods: Similarity Network Fusion (SNF), NEighborhood-based Multi-Omics clustering (NEMO), Data Integration Analysis for Biomarker discovery using Latent variable approaches for Omics studies (DIABLO), and PRAMIGO (Phenotyping netwoRk Application for Multi-omics InteGratiOn), a novel supervised heterogeneous graph transformer (HGT) introduced in this work. Results We benchmarked all five methods, MOFA, DIABLO, SNF, NEMO, and PRAMIGO, on two independent multiomics datasets derived from a Chronic Lymphocytic Leukemia (CLL) cohort comparing IGHV-mutated and unmutated patients, and a pulmonary tuberculosis (PTB) cohort versus healthy controls. Supervised methods (DIABLO, PRAMIGO) consistently achieved higher condition-specific discrimination as measured by the Adjusted Rand Index (ARI) and the Adjusted Mutual Information (AMI). In contrast, unsupervised methods (SNF, NEMO) revealed alternative patient stratifications driven by independent sources of biological variance while MOFA performed in a semi-supervised way occupies an intermediate position, capturing latent factors that explain both disease-associated and orthogonal sources of variance. Systematic gene-centric analysis of the top-ranked features prioritized by each method was supported by manual biological annotation of shared and method-specific signals. Across cohorts, we annotated 154 shared features (116 genes, 38 metabolites) and 60 method-unique features per cohort, demonstrating that no single integration strategy captures the full landscape of biologically relevant signals. In the CLL cohort, shared features spanned B-cell receptor biology, innate immune signaling, RAS/MAPK activation, and epigenetic regulation, while methodunique features revealed supervised-method-specific insights into vesicle trafficking (DIABLO), immune checkpoints (MOFA), and ncRNA regulation (PRAMIGO). In the PTB cohort, a convergent interferon/innate immune signature dominated shared features across all methods. Still, method-unique analysis uncovered DIABLO-specific acylcarnitine metabolic reprogramming, MOFA-specific restoration of lysosomal trafficking, and PRAMIGO-specific {gamma}{delta} T-cell and immunoglobulin repertoire diversity. Across both cohorts, SNF and NEMO proved useful for detecting biological and technical sources of variation that were orthogonal to the primary condition of interest. Ultimately, PRAMIGO uniquely enables the construction of heterogeneous graphs modeling cross-modal molecular interactions, uncovering epigenetic co-regulation programs in CLL and multi-omics inflammatory modules in PTB that are difficult to identify using conventional integration approaches. Conclusions BiomiX 3.0 provides a graphical user interface (GUI) multi-method integration environment that democratizes access to state-of-the-art multi-omics analysis. By combining both unsupervised and supervised integration strategies within a unified platform and introducing graph-based learning through PRAMIGO, BiomiX 3.0 enables researchers across disciplines with complementary tools to interrogate the biological sources of variation in their data, without requiring bioinformatics expertise.

bioinformatics↗

Perfect 21-nucleotide matches to beneficial fungi are common in canonical antifungal dsRNA targets: an in-silico off-target hazard screen for spray-induced gene silencing

Double-stranded RNA (dsRNA) biopesticides that silence essential fungal genes by spray-induced gene silencing (SIGS) are advancing towards registration, underpinned by the claim that silencing is confined to sequence-matched target organisms, a claim that has never been tested systematically against beneficial fungi. We screened twelve dsRNA constructs (the whole transcripts of eleven canonical antifungal target genes of Fusarium graminearum: three CYP51 sterol 14-demethylases, six chitin synthases and two {beta}-tubulins, plus a concatenation of the three CYP51 transcripts modelling the flagship CYP3RNA design) for perfect 21-nucleotide (21-nt) identity, the primary match criterion of recent regulatory bioinformatics frameworks, against the transcriptomes of seven non-target fungi, including commercial biocontrol agents (Trichoderma harzianum, T. virens, Beauveria bassiana, Metarhizium anisopliae, M. brunneum), the arbuscular mycorrhizal fungus Rhizophagus irregularis and Saccharomyces cerevisiae. Eleven of the twelve constructs carried off-target hazard; only CYP51C was completely clean. TUB2b carried 580 perfect 21-nt matches and contained no hazard-free window of 250 consecutive 21-mer start positions, whereas 26.6-60.0% of screened windows were hazard-free in the designable genes. Hits landed overwhelmingly on orthologues of the targeted gene and concentrated in the Hypocreales; R. irregularis and S. cerevisiae were nearly clean. These are hazard findings, not risk findings: perfect siRNA-length matches to beneficial fungi are common in canonical antifungal dsRNA targets, and current regulatory screening panels contain no fungus that would detect them.

bioinformatics↗

Fitting dynamics is not identifying causal edges: a white-box masked ODE benchmark for trans-omics digital twins of drug action mechanisms

Multi-component drug regimens, with traditional Chinese medicine formulas as the hardest case, act across signalling, transcriptional, proteomic and metabolic layers, and elucidating their mechanisms requires dynamic models that predict molecular trajectories rather than static association networks. Trainable ordinary differential equation (ODE) systems fitted to time-series omics are increasingly used for this purpose; yet their validation remains fit-based: at genomic scale, no ground truth has existed to test whether a good fit implies correct mechanisms. Here we build that ground truth: a white-box masked ODE benchmark on the real trans-omics topology of insulin action in mouse liver (transcriptome GEO GSE166336, proteome ProteomeXchange PXD022728, phosphoproteome PXD022823, metabolome source-publication Tables S2-S3; 2,106 molecular species; 4,912 ground-truth edges), with controllable noise, missingness and sampling budgets. Four instruments quantify identifiability: an oracle-perturbation basin curve, a held-out-layer corruption assay, a saturation audit, and an ideal-budget ceiling test. We find that a static baseline (FD + LASSO) performs at chance (AUROC ~ 0.50, except GE at 0.567); that from-scratch training remains at chance even with noise-free, fully observed, densely sampled data (per-layer AUROC 0.48-0.53); that held out layers act as corruption sinks whose failure decomposes into an information floor, a scale-mismatch amplifier, and an edge-gradient drag, curable only jointly; that tanh saturation silently zeroes entire regulator columns; and that a 12-knockout validation battery decomposes intervention reliability by network distance. We distill these into operational prescriptions. The binding constraint is not fitting but structural identifiability. Benchmark, code and audit tools are planned for open release upon publication.

bioinformatics↗