bioRxiv ScienceSearch

Biology subjects

Norman Warthmann

Publications and source records attributed to Norman Warthmann.

3 recordsLinked to original sources

kWIP: The k-mer Weighted Inner Product, a de novo Estimator of Genetic Similarity

Modern genomics techniques generate overwhelming quantities of data. Extracting population genetic variation demands computationally efficient methods to determine genetic relatedness between individuals or samples in an unbiased manner, preferably de novo. The rapid and unbiased estimation of genetic relatedness has the potential to overcome reference genome bias, to detect mix-ups early, and to verify that biological replicates belong to the same genetic lineage before conclusions are drawn using mislabelled, or misidentified samples.\n\nWe present the k-mer Weighted Inner Product (kWIP), an assembly-, and alignment-free estimator of genetic similarity. kWIP combines a probabilistic data structure with a novel metric, the weighted inner product (WIP), to efficiently calculate pairwise similarity between sequencing runs from their k-mer counts. It produces a distance matrix, which can then be further analysed and visualised. Our method does not require prior knowledge of the underlying genomes and applications include detecting sample identity and mix-up, non-obvious genomic variation, and population structure.\n\nWe show that kWIP can reconstruct the true relatedness between samples from simulated populations. By re-analysing several published datasets we show that our results are consistent with marker-based analyses. kWIP is written in C++, licensed under the GNU GPL, and is available from https://github.com/kdmurray91/kwip.\n\nAuthor SummaryCurrent analysis of the genetic similarity of samples is overly dependent on alignment to reference genomes, which are often unavailable and in any case can introduce bias. We address this limitation by implementing an efficient alignment free sequence comparison algorithm (kWIP). The fast, unbiased analysis kWIP performs should be conducted in preliminary stages of any analysis to verify experimental designs and sample metadata, catching catastrophic errors earlier.\n\nkWIP extends alignment-free sequence comparison methods by operating directly on sequencing reads. kWIP uses an entropy-weighted inner product over k-mers as a estimator of genetic relatedness. We validate kWIP using rigorous simulation experiments. We also demonstrate high sensitivity and accuracy even where there is modest divergence between genomes, and/or when sequencing coverage is low. We show high sensitivity in replicate detection, and faithfully reproduce published reports of population structure and stratification of microbiomes. We provide a reproducible workflow for replicating our validation experiments.\n\nkWIP is an efficient, open source software package. Our software is well documented and cross platform, and tutorial-style workflows are provided for new users.

Bioinformatics

Universal metabarcoding of pico- to mesoplankton reveals seasonal dynamics and a bacterial bloom

Most studies of aquatic plankton focus on either macroscopic or microbial communities, and on either eukaryotes or prokaryotes. This separation is primarily for methodological reasons, but can overlook potential interactions among groups. We tested whether DNA-metabarcoding of unfractionated water samples with universal primers could be used to qualitatively and quantitatively study the temporal dynamics of the total plankton community in a shallow temperate lake. We found significant changes in the relative proportions of normalized sequence reads of eukaryotic and prokaryotic plankton communities over a three-month period in spring. Patterns followed the same trend as plankton estimates using traditional microscopic methods. We characterized the bloom of a conditionally rare bacterial taxon belonging to Arcicella, which rapidly came to dominate the whole lake ecosystem and would have remained unnoticed without metabarcoding. Our data demonstrate the potential of universal DNA-metabarcoding applied to unfractionated samples for providing a more holistic view of plankton communities.

Ecology

High habitat-specificity in fungal communities of an oligo-mesotrophic, temperate lake

Freshwater fungi are a poorly studied paraphyletic group that include a high diversity of phyla. Most studies of aquatic fungal diversity have focussed on single habitats, thus the linkage between habitat heterogeneity and fungal diversity remains largely unexplored. We took 216 samples from 54 locations representing eight different habitats in meso-oligotrophic, temperate Lake Stechlin in northern Germany, including the pelagic and littoral water column, sediments, and biotic substrates. We pyrosequenced with an universal eukaryotic marker within the ribosomal large subunit (LSU) in order to compare fungal diversity, community structure, and species turnover among habitats. Our analysis recovered 1024 fungal OTUs (97% criterion). Diversity was highest in the sediment, biofilms, and benthic samples (293-428 OTUs), intermediate in water and reed samples (36-64 OTUs), and lowest in plankton (8 OTUs) samples. NMDS clustering clearly grouped the eight studied habitats into six clusters, indicating that total diversity was strongly influenced by turnover among habitats. Fungal communities exhibited pronounced changes at the levels of phylum and order along a gradient from littoral to pelagic habitats. The large majority of OTUs could not be classified below the order level due to the lack of aquatic fungal entries in taxonomic databases. Our study provides a first estimate of lake-wide fungal diversity and highlights the important contribution of habitat-specificity to total fungal diversity. This remarkable diversity is probably an underestimate, because most lakes undergo seasonal changes and previous studies have uncovered differences in fungal communities among lakes.

Ecology