bioRxiv Science⌕ Search

Biology subjects

Snipen, L.

Publications and source records attributed to Snipen, L..

5 recordsLinked to original sources

Benchmarking 16S rRNA gene amplicon analysis in high-diversity microbial communities reveals fundamental trade-offs in clustering and denoising

Background: Amplicon sequencing of the 16S rRNA gene is widely used to characterise microbial communities, but the performance of commonly used clustering and denoising pipelines has not been systematically evaluated for highly diverse environmental datasets. We therefore benchmarked four established clustering and denoising pipelines, VSEARCH cluster_size, UNOISE as implemented in VSEARCH, Swarm, and DADA2, using simulated microbial communities with known ground-truth compositions and evaluated how performance varied with species richness, abundance unevenness, sequencing depth, and minimum abundance threshold. Results: Differences among pipelines were small at low species richness but became greater as species richness increased. High average cluster purity did not necessarily correspond to accurate species-level recovery, as species could be split across multiple clusters or recovered incompletely. Among the pipelines, UNOISE may be preferable when high species representation and cluster purity are prioritised, while also showing strong reconstruction of relative abundance profiles, albeit with extensive species splitting and many unclustered reads. DADA2 showed similarly strong reconstruction of relative abundance profiles and less species splitting but represented fewer species and had lower cluster purity at higher richness. Swarm provided a balanced compromise between species representation and splitting, whereas cluster_size may be useful when limiting species splitting is a priority, despite weaker reconstruction of relative abundance profiles. Increasing the minimum abundance threshold reduced species splitting but also reduced the number of clusters and perfect clusters and, at higher thresholds, decreased concordance with ground-truth relative abundance profiles. These effects were more pronounced at lower sequencing depths. Analysis of seafloor sediment samples also showed threshold-dependent sequence loss, including the loss of sequences repeatedly detected across replicate samples. Conclusions: Pipeline performance in highly diverse 16S rRNA amplicon datasets varies with dataset characteristics, evaluation criteria, and parameter settings. Pipeline selection should therefore reflect dataset characteristics and the analytical objective rather than rely on a single measure of performance, and minimum abundance thresholds should be evaluated in relation to sequencing depth rather than applied as fixed defaults. These findings emphasise the importance of reconsidering established analytical practices and transparently reporting bioinformatic settings as 16S rRNA amplicon sequencing is increasingly applied to highly diverse environmental microbial communities.

bioinformatics↗

Rsearch: An R interface to VSEARCH supporting visualization and parameter tuning

Background: We present Rsearch, an R package that integrates the core functionality of VSEARCH into the R environment and extends it with visualization, parameter optimization, and conversion tools for compatibility with other R packages. By making VSEARCH directly accessible in R, Rsearch lowers the barrier for using VSEARCH and integrating it with downstream statistical and ecological analyses. Results: Comparative analysis with DADA2 using mock community data showed that both pipelines produced relative abundance profiles highly correlated with the expected composition. Compared to DADA2, Rsearch identified fewer OTUs, but these were more consistently prevalent across samples. In contrast, DADA2 appeared to overestimate diversity by splitting sequences into an excessive number of OTUs. In terms of computational performance, vs_cluster_unoise implemented in Rsearch was the fastest of all the clustering and denoising methods, while other Rsearch functions showed runtimes comparable to DADA2. In addition, Rsearch provides functions for systematic optimization of trimming and filtering parameters, an important feature for users who may not otherwise have a clear strategy for parameter selection. The package also includes functions to ensure compatibility with other R packages such as phyloseq. Conclusions: Rsearch offers a practical and accessible framework for analysing metabarcoding data within a single analytical environment and is freely available from The Comprehensive R Archive Network, with the development version hosted on GitHub (https://github.com/CassandraHjo/Rsearch).

bioinformatics↗

Tracing of streptococcal strains from infant stool across human body sites links gut-specificity to adhesins

Streptococcal species are human commensals known to colonize multiple body sites. Despite being early gut colonizers, we lack strain-level information about their origin and persistence in the gut. To gain more insight into the habitats of the streptococci present in the infant gut, we did a systematic study where mother-infant pairs were sampled from multiple body sites (stool, oral cavity, vagina, breast milk). We performed whole metagenome sequencing and isolated streptococci from 100 infant stool samples (collected at 10 days of age). To trace the streptococci at the strain level, we designed selective qPCR primers for seven streptococcal strains, and these were later utilized to screen corresponding samples from multiple body sites of the infants and their mothers. We found that two of the strains (one Streptococcus parasanguinis and one Streptococcus vestibularis) were highly prevalent in stool samples, both from infants during their first 2 years of life and from their mothers, indicating that these strains are adapted to the gut environment. Interestingly, another S. parasanguinis strain, closely related to the gut-prevalent strain, showed a completely different prevalence pattern, and was mainly detected in vaginal swabs, breast milk and oral swabs. Comparisons of their genomes revealed major differences in genes encoding adhesins, suggesting that host surface attachment could be a key factor for the observed differences in body site specificity. Together, our extensive tracing of streptococci across body sites of 100 infants and their mothers, provides strain-level information of prevalence patterns and reveals the presence of gut-specific streptococci. ImportanceStreptococci thrive on the mucosal surfaces covering the human body and colonize multiple body sites. To determine the distinct streptococcal composition in each habitat and to evaluate their presence in different habitats, strain-level characterization is crucial. We show that two closely related strains, both isolated from stool, are distributed differently across the human body, with one of them prevalent in the stool samples and the other more prevalent in other samples. This emphasizes the necessity of strain-level analysis for the identification of true colonizers of a habitat.

microbiology↗

Standardising a microbiome pipeline for body fluid identification from complex crime scene stains

BackgroundRecent advances in next-generation sequencing have opened up new possibilities for utilizing the human microbiome in various fields, including forensics. Researchers have capitalized on the site-specific microbial communities found in different parts of the body to identify body fluids from biological evidence. Despite promising results, microbiome-based methods have not yet been fully integrated into forensic practice due to the lack of standardized protocols and systematic testing of methods on forensically relevant samples. Our study addresses critical decisions in establishing these protocols, focusing on bioinformatics choices and the use of machine learning to present microbiome results in court for forensically relevant and challenging samples. ResultsWe propose using Operational Taxonomic Units (OTUs) for read data processing and creating heterogeneous training datasets for training a random forest classifier. Our classifier incorporates six forensically relevant classes: saliva, semen, hand skin, penile skin, urine, and vaginal/menstrual fluid. Across these classes, our classifier achieved a high weighted average F1 score of 0.89. Systematic testing on mixed-source samples and underwear revealed reliable detection of at least one component of the mixture and the identification of vaginal fluid from underwear substrates. Additionally, when investigating the sexually shared microbiome (sexome) of heterosexual couples, our classifier shows promising results for the inference of sexual activity. ConclusionIn our study, we recommend the use of a novel random forest classifier trained on a heterogenous dataset for obtaining predictions from samples mimicking forensic evidence. We also highlight the potential of the sexome for assessing the nature of sexual activities in forensic investigations, while delineating areas that warrant further research. Furthermore, we underscore key considerations when presenting machine learning results for classifying mixed-source samples.

bioinformatics↗

HumGut: A comprehensive Human Gut prokaryotic genomes collection filtered by metagenome data

BackgroundA major bottleneck in the use of metagenome sequencing for human gut microbiome studies has been the lack of a comprehensive genome collection to be used as a reference database. Several recent efforts have been made to re-construct genomes from human gut metagenome data, resulting in a huge increase in the number of relevant genomes. In this work, we aimed to create a collection of the most prevalent healthy human gut prokaryotic genomes, to be used as a reference database, including both MAGs from the human gut and ordinary RefSeq genomes. ResultsWe screened > 5,700 healthy human gut metagenomes for the containment of > 490,000 publicly available prokaryotic genomes sourced from RefSeq and the recently announced UHGG collection. This resulted in a pool of > 379,000 genomes that were subsequently scored and ranked based on their prevalence in the healthy human metagenomes. The genomes were then clustered at subspecies resolution, and cluster representatives were retained to comprise the HumGut collection. Using the Kraken2 software for classification, we find superior performance in the assignment of metagenomic reads, classifying on average 94.5% of the reads in a metagenome, as opposed to 86% with UHGG and 44% when using standard Kraken2 database. HumGut, half the size of standard Kraken2 database and directly comparable to the UHGG size, outperforms them both. ConclusionsThe HumGut collection contains > 30,000 genomes clustered at subspecies resolution and ranked by human gut prevalence. We demonstrate how metagenomes from IBD-patients map equally well to this collection, indicating this reference is relevant also for studies well outside the metagenome reference set used to obtain HumGut. We believe this is a valuable resource in a field in dire need of method standardization. All data and metadata, as well as helpful code, are available at http://arken.nmbu.no/~larssn/humgut/.

microbiology↗