bioRxiv · 10.1101/768713
Machine Learning in Quality Assessment of Early Stage Next-Generation Sequencing Data
Abstract
Controlling quality of next generation sequencing (NGS) data files is a necessary but complex task. To address this problem, we statistically characterized common NGS quality features and developed a novel quality control procedure involving tree-based and deep learning classification algorithms. Predictive models, validated on internal data and external disease diagnostic datasets, are to some extent generalizable to data from unseen species. The derived statistical guidelines and predictive models represent a valuable resource for users of NGS data to better understand quality issues and perform automatic quality control. Our guidelines and software are available at the following URL: https://github.com/salbrec/seqQscorer.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Albrecht, S., Andrade-Navarro, M. A., Fontaine, J.-F.. 2019-09-14. Machine Learning in Quality Assessment of Early Stage Next-Generation Sequencing Data. https://doi.org/10.1101/768713
Cite the original work for its findings. Save a collection to share your selection of sources.