bioRxiv · 10.1101/2020.04.30.066159
Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms
Abstract
PurposeClinical genome sequencing (cGS) followed by orthogonal confirmatory testing is standard practice. While orthogonal testing significantly improves specificity it also results in increased turn-around-time and cost of testing. The purpose of this study is to evaluate machine learning models trained to identify false positive variants in cGS data to reduce the need for orthogonal testing. MethodsWe sequenced five reference human genome samples characterized by the Genome in a Bottle Consortium (GIAB) and compared the results to an established set of variants for each genome referred to as a truth-set. We then trained machine learning models to identify variants that were labeled as false positives. ResultsAfter training, the models identified 99.5% of the false positive heterozygous single nucleotide variants (SNVs) and heterozygous insertions/deletions variants (indels) while reducing confirmatory testing of true positive SNVs to 1.67% and indels to 20.29%. Employing the algorithm in clinical practice reduced orthogonal testing using dideoxynucleotide (Sanger) sequencing by 78.22%. ConclusionOur results indicate that a low false positive call rate can be maintained while significantly reducing the need for confirmatory testing. The framework that generated our models and results is publicly available at https://github.com/HudsonAlpha/STEVE.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Holt, J. M., Wilk, M., Sundlof, B., Nakouzi, G., Bick, D., Lyon, E.. 2020-05-02. Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms. https://doi.org/10.1101/2020.04.30.066159
Cite the original work for its findings. Save a collection to share your selection of sources.