bioRxiv Science⌕ Search

Biology subjects

Yuan, T.-H.

Publications and source records attributed to Yuan, T.-H..

2 recordsLinked to original sources

Toward Automatic Variant Interpretation: Discordant Genetic Interpretation Across Variant Annotations for ClinVar Pathogenic Variants

PurposeHigh-throughput sequencing has revolutionized genetic disorder diagnosis, but variant pathogenicity interpretation is still challenging. Even though the Human Genome Variation Society (HGVS) provides recommendations for variant nomenclature, discrepancies in annotation remain a significant hurdle. MethodsThis study evaluated the annotation concordance between three tools-- ANNOVAR, SnpEff, and Variant Effect Predictor (VEP)--using 164,549 two-star variants from ClinVar. The analysis used HGVS nomenclature string-match comparisons to assess annotation consistency from each tool, corresponding coding impacts, and associated ACMG criteria inferred from the annotations. ResultsThe analysis revealed variable concordance rates, with 58.52% agreement for HGVSc, 84.04% for HGVSp, and 85.58% for the coding impact. SnpEff showed the highest match for HGVSc (0.988), while VEP bettered for HGVSp (0.977). The substantial discrepancies were noted in the Loss-of-Function (LoF) category. Incorrect PVS1 interpretations affected the final pathogenicity and downgraded PLP variants (ANNOVAR 55.9%, SnpEff 66.5%, VEP 67.3%), risking false negatives of clinically relevant variants in reports. ConclusionsThese findings highlight the critical challenges in accurately interpreting variant pathogenicity due to discrepancies in annotations. To enhance the reliability of genetic variant interpretation in clinical practice, standardizing transcript sets and systematically cross-validating results across multiple annotation tools is essential. Graphic abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=120 SRC="FIGDIR/small/617756v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@9f6378org.highwire.dtl.DTLVardef@3b9a2corg.highwire.dtl.DTLVardef@106fb58org.highwire.dtl.DTLVardef@15f70c8_HPS_FORMAT_FIGEXP M_FIG This study examined the consistency of variant annotations produced by three widely used open-source toolsANNOVAR, SnpEff, and VEPagainst 164,549 ClinVar two starts variants. The investigation covers HGVS-based transcript, protein nomenclature and coding impact annotation. The results showed that none of the tools were fully consistent with ClinVar across all coding impact categories, particularly in the LoF category, which exhibited the poorest consistency. This inconsistency may lead to discrepancies in PVS1 interpretation, affecting the final pathogenicity assessment. PVS1 loss resulted in a significant downgrading of PLP variants, potentially leading to the omission of clinically relevant variants in reports. C_FIG

genomics↗

Somatic mutation detection workflow validity distinctly influences clinical decision.

Evaluating robustness of somatic mutation detections is essential when utilizing whole exome sequencing (WES) for treatment decision-making. A comprehensive evaluation was conducted using tumor WES from the FDA-led Sequencing Quality Control Phase 2 (SEQC2) project, in which multiple library kits sequenced identical DNA materials across three labs to benchmark analytical validity. These workflows included various read aligner (BWA, Bowtie2, DRAGEN-Aligner, DRAGMAP, and HISAT2) and mutation caller (Mutect2, TNscope, DRAGEN-Caller, and DeepVariant) combinations. The results revealed that DRAGEN exhibited superior performance, achieving mean F1-scores of 0.966 and 0.791 for SNV and INDEL detection, respectively. Among open-source software, BWA Mutect2 and HISAT2 Mutect2 combinations showed the highest mean F1-scores for SNV (0.949) and IN-DEL (0.722), respectively. The analyses indicated that high-quality data can be analyzed as having worse results, and vice versa. Evaluations of COSMIC reported mutations unveiled discrepancies across enrichment kits. IDT enrichment kits showed a higher false negative rate, while Agilent WES kits tended to miss mutations in CBL and IDH1, and Roche library kits tended to miss the mutations in PIK3CB. For drug-related biomarkers, Sentieon TNscope tended to underestimate tumor mutation burden and overlook crucial drug-resistance mutations such as FLT3 (c.G1879A: p.A627T) for cytarabine resistance in leukemia and MAP2K1 (c.G199A:p.D67N) for BRAF inhibitors in melanoma. The findings highlight the importance of robust bioinformatic analysis in identifying tumor mutations and guiding clinical decision-making. HighlightsO_LIMutation callers had a significantly higher effect on overall sensitivity than aligners. C_LIO_LIBenchmarking analyses demonstrated that high-quality sequencing reads can be analyzed as having worse results, and vice versa. C_LIO_LIDRAGEN exhibited the best performance among other aligner-caller combinations. C_LIO_LIThe combination of BWA with Mutect2 and HISAT2 with Mutect2 yielded the highest mean F1 scores for detecting SNVs and INDELs by open-source software, respectively. C_LIO_LISentieon TNscope tended to underestimate the tumor mutation burden and missed several drug-resistant mutations. C_LI

bioinformatics↗