bioRxiv ScienceSearch

Biology subjects

Salzberg, S. L.

Publications and source records attributed to Salzberg, S. L..

3 recordsLinked to original sources

KrakenHLL: Confident and fast metagenomics classification using unique k-mer counts

False positive identifications are a significant problem in metagenomic classification. We present KrakenHLL, a novel metagenomic classifier that combines the fast k-mer based classification of Kraken with an efficient algorithm for assessing the coverage of unique k-mers found in each species in a dataset. On various test datasets, KrakenHLL gives better recall and F1-scores than other methods, and effectively classifies and distinguishes pathogens with low abundance from false positives in infectious disease samples. By using the probabilistic cardinality estimator HyperLogLog (HLL), KrakenHLL is as fast as Kraken and requires little additional memory.

bioinformatics

The first near-complete assembly of the hexaploid bread wheat genome, Triticum aestivum

Common bread wheat, Triticum aestivum, has one of the most complex genomes known to science, with 6 copies of each chromosome, enormous numbers of near-identical sequences scattered throughout, and an overall size of more than 15 billion bases. Multiple past attempts to assemble the genome have failed. Here we report the first successful assembly of T. aestivum, using deep sequencing coverage from a combination of short Illumina reads and very long Pacific Biosciences reads. The final assembly contains 15,344,693,583 bases and has a weighted average (N50) contig size of of 232,659 bases. This represents by far the most complete and contiguous assembly of the wheat genome to date, providing a strong foundation for future genetic studies of this important food crop. We also report how we used the recently published genome of Aegilops tauschii, the diploid ancestor of the wheat D genome, to identify 4,179,762,575 bp of T. aestivum that correspond to its D genome components.

genomics

Pavian: Interactive analysis of metagenomics data for microbiomics and pathogen identification

SummaryPavian is a web application for exploring metagenomics classification results, with a special focus on infectious disease diagnosis. Pinpointing pathogens in metagenomics classification results is often complicated by host and laboratory contaminants as well as many non-pathogenic microbiota. With Pavian, researchers can analyze, display and transform results from the Kraken and Centrifuge classifiers using interactive tables, heatmaps and flow diagrams. Pavian also provides an alignment viewer for validation of matches to a particular genome.\n\nAvailability and implementationPavian is implemented in the R language and based on the Shiny framework. It can be hosted on Windows, Mac OS X and Linux systems, and used with any contemporary web browser. It is freely available under a GPL-3 license from http://github.com/fbreitwieser/pavian. Furthermore a Docker image is provided at https://hub.docker.com/r/florianbw/pavian.\n\nContactfbreitw1@jhu.edu\n\nSupplementary informationSupplementary data is available at Bioinformatics online.

bioinformatics