bioRxiv ScienceSearch

Biology subjects

Koslicki, D.

Publications and source records attributed to Koslicki, D..

7 recordsLinked to original sources

Assessing taxonomic metagenome profilers with OPAL

Taxonomic metagenome profilers predict the presence and relative abundance of microorganisms from shotgun sequence samples of DNA isolated directly from a microbial community. Over the past years, there has been an explosive growth of software and algorithms for this task, resulting in a need for more systematic comparisons of these methods based on relevant performance criteria. Here, we present OPAL, a software package implementing commonly used performance metrics, including those of the first challenge of the Initiative for the Critical Assessment of Metagenome Interpretation (CAMI), together with convenient visualizations. In addition, OPAL implements diversity metrics from microbial ecology, as well as run time and memory efficiency measurements. By allowing users to customize the relative importance of metrics, OPAL facilitates in-depth performance comparisons, as well as the development of new methods and data analysis workflows. To demonstrate the application, we compared seven profilers on benchmark datasets of the first and second CAMI challenges using all metrics and performance measurements available in OPAL. The software is implemented in Python 3 and available under the Apache 2.0 license on GitHub (https://github.com/CAMI-challenge/OPAL).\n\nAuthor summaryThere are many computational approaches for inferring the presence and relative abundance of taxa (i.e. taxonomic profiling) from shotgun metagenome samples of microbial communities, making systematic performance evaluations a very important task. However, there has yet to be introduced a computational framework in which profiler performances can be compared. This delays method development and applied studies, as researchers need to implement their own custom evaluation frameworks. Here, we present OPAL, a software package that facilitates standardized comparisons of taxonomic metagenome profilers. It implements a variety of performance metrics frequently employed in microbiome research, including runtime and memory usage, and generates comparison reports and visualizations. OPAL thus facilitates and accelerates benchmarking of taxonomic profiling techniques on ground truth data. This enables researchers to arrive at informed decisions about which computational techniques to use for specific datasets and research questions.

bioinformatics

MiCoP: Microbial Community Profiling method for detecting viral and fungal organisms in metagenomic samples

BackgroundHigh throughput sequencing has spurred the development of metagenomics, which involves the direct analysis of microbial communities in various environments such as soil, ocean water, and the human body. Many existing methods based on marker genes or k-mers have limited sensitivity or are too computationally demanding for many users. Additionally, most work in metagenomics has focused on bacteria and archaea, neglecting to study other key microbes such as viruses and eukaryotes.\n\nResultsHere we present a method, MiCoP (Microbiome Community Profiling), that uses fast-mapping of reads to build a comprehensive reference database of full genomes from viruses and eukaryotes to achieve maximum read usage and enable the analysis of the virome and eukaryome in each sample. We demonstrate that mapping of metagenomic reads is feasible for the smaller viral and eukaryotic reference databases. We show that our method is accurate on simulated and mock community data and identifies many more viral and fungal species than previously-reported results on real data from the Human Microbiome Project.\n\nConclusionsMiCoP is a mapping-based method that proves more effective than existing methods at abundance profiling of viruses and eukaryotes in metagenomic samples. MiCoP can be used to detect the full diversity of these communities. The code, data, and documentation is publicly available on GitHub at: https://github.com/smangul1/MiCoP

bioinformatics

Improving Min Hash via the Containment Index with applications to Metagenomic Analysis

Min hash is a probabilistic method for estimating the similarity of two sets in terms of their Jaccard index, defined as the ration of the size of their intersection to their union. We demonstrate that this method performs best when the sets under consideration are of similar size and the performance degrades considerably when the sets are of very different size. We introduce a new and efficient approach, called the containment min hash approach, that is more suitable for estimating the Jaccard index of sets of very different size. We accomplish this by leveraging another probabilistic method (in particular, Bloom filters) for fast membership queries. We derive bounds on the probability of estimate errors for the containment min hash approach and show it significantly improves upon the classical min hash approach. We also show significant improvements in terms of time and space complexity. As an application, we use this method to detect the presence/absence of organisms in a metagenomic data set, showing that it can detect the presence of very small, low abundance microorganisms.

bioinformatics

IndeCut evaluates performance of network motif discovery algorithms

Genomic networks represent a complex map of molecular interactions which are descriptive of the biological processes occurring in living cells. Identifying the small over-represented circuitry patterns in these networks helps generate hypotheses about the functional basis of such complex processes. Network motif discovery is a systematic way of achieving this goal. However, a reliable network motif discovery outcome requires generating random background networks which are the result of a uniform and independent graph sampling method. To date, there has been no sound practical method to numerically evaluate whether any network motif discovery algorithm performs as intended--thus it was not possible to assess the validity of resulting network motifs. In this work, we present IndeCut, the first and only method that allows characterization of network motif finding algorithm performance on any network of interest. We demonstrate that it is critical to use IndeCut prior to running any network motif finder for two reasons. First, IndeCut estimates the minimally required number of samples that each network motif discovery tool needs in order to produce an outcome that is both reproducible and accurate. Second, IndeCut allows users to choose the most accurate network motif discovery tool for their network of interest among many available options. IndeCut is an open source software package and is available at https://github.com/megrawlab/IndeCut.

systems biology

Critical Assessment of Metagenome Interpretation - a benchmark of computational metagenomics software

In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.

bioinformatics

EMDUnifrac: Exact linear time computation of the Unifrac metric and identification of differentially abundant organisms

Both the weighted and unweighted Unifrac distances have been very successfully employed to assess if two communities differ, but do not give any information about how two communities differ. We take advantage of recent observations that the Unifrac metric is equivalent to the so-called earth movers distance (also known as the Kantorovich-Rubinstein metric) to develop an algorithm that not only computes the Unifrac distance in linear time and space, but also simultaneously finds which operational taxonomic units are responsible for the observed differences between samples. This allows the algorithm, called EMDUnifrac, to determine why given samples are different, not just if they are different, and with no added computational burden. EMDUnifrac can be utilized on any distribution on a tree, and so is particularly suitable to analyzing both operational taxonomic units derived from amplicon sequencing, as well as community profiles resulting from classifying whole genome shotgun metagenomes. The EMDUnifrac source code (written in python) is freely available at: https://github.com/dkoslicki/EMDUnifrac.

microbiology

Exact probabilities for the indeterminacy of complex networks as perceived through press perturbations

We consider the goal of predicting how complex networks respond to chronic (press) perturbations when characterizations of their network topology and interaction strengths are associated with uncertainty. Our primary result is the derivation of exact formulas for the expected number and probability of qualitatively incorrect predictions about a systems responses under uncertainties drawn form arbitrary distributions of error. These formulas obviate the current use of simulations, algorithms, and qualitative modeling techniques. Additional indices provide new tools for identifying which links in a network are most qualitatively and quantitatively sensitive to error, and for determining the volume of errors within which predictions will remain qualitatively determinate (i.e. sign insensitive). Together with recent advances in the empirical characterization of uncertainty in ecological networks, these tools bridge a way towards probabilistic predictions of network dynamics.

ecology