bioRxiv Science⌕ Search

Biology subjects

Monger, S.

Publications and source records attributed to Monger, S..

2 recordsLinked to original sources

Benchmarking of variant pathogenicity prediction methods using a population genetics approach

MotivationVariant pathogenicity predictors are essential for identifying new associations between genetic variants and rare diseases. However, despite the availability of numerous predictors, there is no clear consensus on which methods provide the most reliable results. The common practice of training, testing, and benchmarking these predictors using known variant sets from disease or mutagenesis studies raises concerns about ascertainment bias and data circularity. ResultsWe benchmarked commonly used pathogenicity predictors using an orthogonal approach that does not rely on predefined "ground truth" datasets. By leveraging population-level genomic data from gnomAD and the Context-Adjusted Proportion of Singletons (CAPS) metric, we identified CADD and REVEL as the best-performing predictors for distinguishing extremely deleterious variants from moderately deleterious ones. REVEL demonstrated superior calibration. Additionally, we show that CAPS can serve as a meta-analysis tool for interpreting variant annotations and highlight biases in ClinVar-based predictor training. Availability and ImplementationCAPS analysis and benchmarking results are available at https://github.com/mgudVCCRI/PopGenVariantFiltering Contacte.giannoulatou@victorchang.edu.au

genomics↗

Systematic evaluation of de novo mutation calling tools using whole genome sequencing data

De novo mutations (DNMs) are genetic alterations that occur for the first time in an offspring. DNMs have been found to be a significant cause of severe developmental disorders. With the widespread use of next-generation sequencing (NGS) technologies, accurate detection of DNMs is crucial. Several bioinformatics tools have been developed to call DNMs from NGS data, but no study to date has systematically compared these tools. We used both real whole genome sequencing (WGS) data from a trio from the 1000 Genomes Project (1000G) and an in-house simulated trio dataset to evaluate five DNM calling tools: DeNovoGear, TrioDeNovo, PhaseByTransmission, VarScan2, and DeNovoCNN. For DNMs called in the real dataset, we observed 8.4% concordance of variants between all tools, while 83.8% of DNMs variants were identified by only one caller. For simulated trio WGS dataset spiked with 100 DNMs, the concordance rate was also low at 3.9%. DeNovoGear achieved the highest F1 score on the real 1000G dataset, while DeNovoCNN had the highest F1 score on the simulated data. Our study provides valuable recommendations for the selection and application of DNM callers on WGS trio data.

bioinformatics↗