bioRxiv · 10.1101/237107
Cleaning clinical genomic data: Simple identification and removal of recurrently miscalled variants in single genomes
Abstract
Identification of sequence variation from short-read sequence data is subject to common-yet-intermittent miscalling that occurs in a sequence intrinsic manner. We identify that recurrent false positive single nucleotide variants are strongly present in databases of human sequence variation and demonstrate how each individual sample generates a unique set of recurrent false positive variants. These recurrent miscalls result from known difficulties aligning short-read sequence data between redundant genomic regions. We could replicate, catalogue and remove three quarters of these recurrent miscalls for any given exome with as little as ten rounds of read resampling, realignment and recalling. The removal of such misleading variants reduces the search space for identification of disease causing variants.\n\nList of Abbreviations
Explore related subjects
Keep this discovery
Field, M. A., Burgio, G., Al Shekaili, J., Foote, S. J., Cook, M. C., Andrews, T. D.. 2017-12-20. Cleaning clinical genomic data: Simple identification and removal of recurrently miscalled variants in single genomes. https://doi.org/10.1101/237107
Cite the original work for its findings. Save a collection to share your selection of sources.