bioRxiv Science⌕ Search

Biology subjects

Kralj, J. G.

Publications and source records attributed to Kralj, J. G..

5 recordsLinked to original sources

A mathematical framework to correct for compositionality in microbiome datasets

The increasing use of metagenomic sequencing (MGS) for microbiome analysis has significantly advanced our understanding of microbial communities and their roles in various biological processes, including human health, environmental cycling, and disease. However, the inherent compositionality of MGS data, where the relative abundance of each taxa depends on the abundance of all other taxa, complicates the measurement of individual taxa and the interpretation of microbiome data. Here we describe an experimental design that incorporates exogenous internal standards in routine MGS analyses to correct for compositional distortions. A mathematical framework was developed for using the observed internal standard relative abundance to calculate "Scaled Abundances" for native taxa that were (i) independent of sample composition and (ii) directly proportional to actual biological abundances. Through rigorous analysis of mock community and human gut microbiome samples, we demonstrate that Scaled Abundances outperformed traditional relative abundance measurements in both precision and accuracy and enabled reliable, quantitative comparisons of individual microbiome taxa across varied sample compositions and across a wide range of taxa abundances. By providing a pathway to accurate taxa quantification, this approach holds significant potential for advancing microbiome research, particularly in clinical and environmental health applications where precise microbial profiling is critical. ImportanceMetagenomic sequencing (MGS) analysis has become central to modern characterizations of microbiome samples. However, the inherent compositionality of these analyses often complicate interpretations of results. We present here an experimental design and corresponding mathematical framework that uses internal standards with routine MGS methods to correct for compositional distortions. We validate this approach for both amplicon and shotgun MGS analysis of mock communities and human gut microbiome (fecal) samples. By using internal standards to remove compositionality, we demonstrate significantly improved measurement accuracy and precision for quantification of taxa abundances. This approach is broadly applicable across a wide range of microbiome research applications.

microbiology↗

Analytical Assessment of Metagenomic Workflows for Pathogen Detection with NIST RM 8376 and Two Sample Matrices

We assessed the analytical performance of metagenomic workflows using NIST Reference Material 8376 DNA from bacterial pathogens spiked into two simulated clinical samples: cerebral spinal fluid (CSF) and stool. Sequencing and taxonomic classification were used to generate signals for each sample and taxa of interest, and to estimate the limit of detection (LOD), the response function, and linear dynamic range. We found that the LODs for taxa spiked into CSF ranged from approximately (0.1 to 0.3) copy/L, with a linearity of 0.96 to 0.99. For stool, the LODs ranged from (10 to 221) copy/L, with a linearity of 0.99 to 1.01. Further, discriminating different E. coli strains proved to be workflow-dependent, as only one classifier:database combination of the three tested showed the ability to differentiate the two pathogenic and commensal strains. Surprisingly, when we compared the response functions of the same taxa in the two different sample types, we found those functions to be the same, despite large differences in LODs. This suggests that the "agnostic diagnostic" theory for metagenomics may apply to different target organisms and different sample types. Using RMs, we were able to generate quantitative analytical performance metrics for each workflow and sample set, enabling relatively rapid workflow screening before employing clinical samples. This makes these RMs a useful tool that will generate data needed to support translation of metagenomics into regulated use. ImportanceAssessing the analytical performance of metagenomic workflows, especially when developing clinical diagnostics, is foundational for ensuring that the measurements underlying a diagnosis are supported by rigorous characterization. To facilitate the translation of metagenomics into clinical practice, workflows must be tested using control samples designed to probe the analytical limitations (e.g. limit of detection). Spike-ins allow developers to generate fit-for-purpose control samples for initial workflow assessments and inform decisions about further development. However, clinical sample types include a wide range of compositions and concentrations, each presenting different detection challenges. In this work, we demonstrate how spike-ins elucidate workflow performance in two highly dissimilar sample types (stool and CSF); and we provide evidence that detection of individual organisms is unaffected by background sample composition, making detection sample agnostic within a workflow. These demonstrations and performance insights will facilitate translation of the technology to the clinic.

bioinformatics↗

A Sensitivity Analysis of Methodological Variables Associated with Microbiome Measurements

The experimental methods employed during metagenomic sequencing analyses of microbiome samples significantly impact the resulting data and typically vary substantially between laboratories. In this study, a full factorial experimental design was used to compare the effects of a select set of methodological choices (sample, operator, lot, extraction kit, variable region, reference database) on the analysis of biologically diverse stool samples. For each parameter investigated, a main effect was calculated that allowed direct comparison both between methodological choices (bias effects) and between samples (real biological differences). Overall, methodological bias was found to be similar in magnitude to real biological differences, while also exhibiting significant variations between individual taxa, even between closely related genera. The quantified method biases were then used to computationally improve the comparability of datasets collected under substantially different protocols. This investigation demonstrates a framework for quantitatively assessing methodological choices that could be routinely performed by individual laboratories to better understand their metagenomic sequencing workflows and to improve the scope of the datasets they produce.

microbiology↗

Variability and Bias in Microbiome Metagenomic Sequencing: an Interlaboratory Study Comparing Experimental Protocols

BackgroundSeveral studies have documented the significant impact of methodological choices in microbiome analyses. The myriad of methodological options available complicate the replication of results and generally limit the comparability of findings between independent studies that use differing techniques and measurement pipelines. Here we describe the Mosaic Standards Challenge (MSC), an international interlaboratory study designed to assess the impact of methodological variables on the results. The MSC did not prescribe methods but rather asked participating labs to analyze 7 shared reference samples (5x human stool samples and 2x mock communities) using their standard laboratory methods. To capture the array of methodological variables, each participating lab completed a metadata reporting sheet that included 100 different questions regarding the details of their protocol. The goal of this study was to survey the methodological landscape for microbiome metagenomic sequencing (MGS) analyses and the impact of methodological decisions on metagenomic sequencing results. ResultsA total of 44 labs participated in the MSC by submitting results (16S or WGS) along with accompanying metadata; thirty 16S rRNA gene amplicon datasets and 14 WGS datasets were collected. The inclusion of two types of reference materials (human stool and mock communities) enabled analysis of both MGS measurement variability between different protocols using the biologically-relevant stool samples, and MGS bias with respect to ground truth values using the DNA mixtures. Owing to the compositional nature of MGS measurements, analyses were conducted on the ratio of Firmicutes: Bacteroidetes allowing us to directly apply common statistical methods. The resulting analysis demonstrated that protocol choices have significant effects, including both bias of the MGS measurement associated with a particular methodological choices, as well as effects on measurement robustness as observed through the spread of results between labs making similar methodological choices. In the analysis of the DNA mock communities, MGS measurement bias was observed even when there was general consensus among the participating laboratories. ConclusionThis study was the result of a collaborative effort that included academic, commercial, and government labs. In addition to highlighting the impact of different methodological decisions on MGS result comparability, this work also provides insights for consideration in future microbiome measurement study design.

microbiology↗

Considerations for performance metrics of metagenomic next generation sequencing analyses

Evaluating the performance of metagenomics analyses has proven a challenge, due in part to limited ground-truth standards, broad application space, and numerous evaluation methods and metrics. Application of traditional clinical performance metrics (i.e. sensitivity, specificity, etc.) using taxonomic classifiers do not fit the "one-bug-one-test" paradigm. Ultimately, users need methods that evaluate fitness-for-purpose and identify their analyses strengths and weaknesses. Within a defined cohort, reporting performance metrics by taxon, rather than by sample, will clarify this evaluation. An estimated limit of detection, positive and negative control samples, and true positive and negative true results are necessary criteria for all investigated taxa. Use of summary metrics should be restricted to comparing results of similar cohorts and data, and should employ harmonic means and continuous products for each performance metric rather than arithmetic mean. Such consideration will ensure meaningful comparisons and evaluation of fitness-for-purpose.

genomics↗