bioRxiv Science⌕ Search

Biology subjects

Lee, D.-e.

Publications and source records attributed to Lee, D.-e..

2 recordsLinked to original sources

Quantitative Optimization of Sensitivity and Specificity in Targeted and Whole-Exome Sequencing Using Reference-Standard DNA Mixtures

BackgroundWe previously developed a benchmarking strategy using mixtures of homozygote and heterozygote DNAs as reference standards to simultaneously assess sensitivity and false positive (FP) error rates in targeted next-generation sequencing (T-NGS) and whole-exome sequencing (WES), revealing substantial variability across commercial platforms. However, optimal analytic conditions for clinical application remain undefined. MethodsWe systematically evaluated multiple sequencing kits and bioinformatics pipelines across various variant allele fraction (VAF) thresholds to identify conditions that maximize both sensitivity and specificity. Recurrent error-prone alleles were defined and filtered to enhance specificity. ResultsOptimal performance was achieved using the DRAGEN pipeline with recurrent FP allele filtering. For T-NGS, a 1% VAF cutoff yielded a 95% detection threshold of 2.99% and 1.21 FPs per megabase (FP/Mb); for WES, a 2% cutoff yielded a 95% threshold of 5.02% and 1.15 FP/Mb. These settings improved sensitivity >3-fold and reduced FP rates >96% versus suboptimal pipelines. Notably, VAF thresholds flattened sensitivity differences across platforms, obscuring key performance disparities--challenging assumptions that T-NGS is inherently more sensitive than WES. In-house and conventional pipelines undercalled up to 10% of true variants. Restricting reporting of 1-4% VAF variants to [~]1,000 predefined actionable sites enabled recovery of clinically relevant mutations while reducing FP risk >99%. ConclusionsThis study provides a quantitative framework for optimizing NGS performance. Our findings support actionable strategies to improve diagnostic accuracy in clinical genomics through tailored pipeline selection, VAF thresholding, and artifact filtering.

bioinformatics↗

Evaluation of false positive and false negative errors in targeted next generation sequencing

BackgroundAlthough next generation sequencing (NGS) has been adopted as an essential diagnostic tool in various diseases, NGS errors have been the most serious problem in clinical implementation. Especially in cancers, low level mutations have not been easy to analyze, due to the contaminating normal cells and tumor heterozygosity. ResultsIn targeted NGS (T-NGS) analyses for reference-standard samples containing mixtures of homozygote H. mole DNA with blood genomic DNA at various ratios from four certified NGS service providers, large differences in the lower detection limit of variants (16.3 times, 1.51[~]24.66%) and the false positive (FP) error rate (4280 times, 5.814 x 10-4 [~]1.359 x 10-7) were found. Employment of the commercially available Dragen system for bioinformatic analyses reduced FP errors in the results from companies BB and CC, but the errors originating from the NGS raw data persisted. Bioinformatic conditional adjustment to increase sensitivity (less than 2 times) led to a much higher FP error rate (610[~]8200 times). In addition, problems such as biased preferential reference base calls during bioinformatic analysis and high-rate FN errors in HLA regions were found in the NGS analysis. ConclusionT-NGS results from certified NGS service providers can be quite various in their sensitivity and FP error rate, suggesting the necessity of further quality controls for clinical implementation of T-NGS. The present study also suggests that mixtures of homozygote and heterozygote DNAs can be easily employed as excellent reference-standard materials for quality control of T-NGS.

genomics↗