bioRxiv ScienceSearch

Biology subjects

Emrich, S.

Publications and source records attributed to Emrich, S..

3 recordsLinked to original sources

The Effects of Normalization, Transformation, and Rarefaction on Clustering of OTU Abundance

Introduction Introduction Methods Discussion Data and Code Availability References Metagenomic clustering presents a unique opportunity to associate and understand communities. Working with Operational Taxonomic Units (OTUs), however, often requires a strategy for handling OTUs that may be over or under represented in a given sample, which is thought of as \"erroneous\". PCR amplification, for example, is known to sometimes non-linearly over-represent more common species Gonzalez et al. 2012. Strategies dealing with this bias include Normalization, Rarefaction, and Log Transformation. Here, we examine how methods to handle potential outlier observations affect de novo estimation of groups using both clustering and matrix factorization methods.\n\nWhile log transformations may affect parametric tests ...

ecology

Methods in Description and Validation of Local Metagenetic Microbial Communities

1. We propose MinHash (as implemented by MASH) and NMF as alternative methods to estimate similarity between metagenetic samples. We further describe these results with cluster analysis and correlations with independent ecological metadata.\n\n2. Using sample to sample similarities based on MinHash similarities we use hierarchal clustering to generate clusters, simultaneously we generate groups based on NMF, and we compare groups generated from the MinHash similarity derived clusters and from NMF to those determined by the environment, looking to Silhouette Width for an assessment of the quality of the cluster.\n\n3. We analyze existing data from the Atacama Desert to determine the relationship between ecological factors and group membership, and using the generated groups from MASH and NMF we run an ANOVA to uncover links between metagenetic samples and known environmental variables such as pH and Soil Conductivity.

bioinformatics

HECIL: A Hybrid Error Correction Algorithm for Long Reads with Iterative Learning

Second-generation sequencing techniques generate short reads that can result in fragmented genome assemblies. Third-generation sequencing platforms mitigate this limitation by producing longer reads that span across complex and repetitive regions. Currently, the usefulness of such long reads is limited, however, because of high sequencing error rates. To exploit the full potential of these longer reads, it is imperative to correct the underlying errors. We propose HECIL--Hybrid Error Correction with Iterative Learning--a hybrid error correction framework that determines a correction policy for erroneous long reads, based on optimal combinations of decision weights obtained from short read alignments. We demonstrate that HECIL outperforms state-of-the-art error correction algorithms for an overwhelming majority of evaluation metrics on diverse real data sets including E. coli, S. cerevisiae, and the malaria vector mosquito A. funestus. We further improve the performance of HECIL by introducing an iterative learning paradigm that improves the correction policy at each iteration by incorporating knowledge gathered from previous iterations via confidence metrics assigned to prior corrections.\n\nAvailability and Implementationhttps://github.com/NDBL/HECIL\n\nContactsemrich@nd.edu

bioinformatics