bioRxiv Science⌕ Search

Biology subjects

Jaspez, D.

Publications and source records attributed to Jaspez, D..

3 recordsLinked to original sources

A benchmarking of human Y-chromosomal haplogroup classifiers from whole-genome and whole-exome sequence data

The non-recombinant region of the Y chromosome (NRY) contains a great number of polymorphic markers that allows to accurately reconstruct pedigree relationships and retrieve ancestral information from study samples. The analysis of NRY is typically implemented in anthropological, medical, and forensic studies. High-throughput sequencing (HTS) has profoundly increased the identification of genetic markers in the NRY genealogy and has prompted the development of automated NRY haplogroup classification tools. Here, we present a benchmarking study of five command-line tools for NRY haplogroup classification. The evaluation was done using empirical short-read HTS data from 50 unrelated donors using paired data from whole-genome sequencing (WGS) and whole-exome sequencing (WES) experiments. Besides, we evaluate the performance of the top-ranked tool in the classification of data of third generation HTS obtained from a subset of donors. Our findings demonstrate that WES can be an efficient approach to infer the NRY haplogroup, albeit generally providing a lower level of genealogical resolution than that recovered by WGS. Among the tools evaluated, YLeaf offers the best performance for both WGS and WES applications. Finally, we demonstrate that YLeaf is able to correctly classify all samples sequenced with nanopore technology from long noisy reads.

genomics↗

Towards a Comprehensive Variation Benchmark for Challenging Medically-Relevant Autosomal Genes

The repetitive nature and complexity of multiple medically important genes make them intractable to accurate analysis, despite the maturity of short-read sequencing, resulting in a gap in clinical applications of genome sequencing. The Genome in a Bottle Consortium has provided benchmark variant sets, but these excluded some medically relevant genes due to their repetitiveness or polymorphic complexity. In this study, we characterize 273 of these 395 challenging autosomal genes that have multiple implications for medical sequencing. This extended, curated benchmark reports over 17,000 SNVs, 3,600 INDELs, and 200 SVs each for GRCh37 and GRCh38 across HG002. We show that false duplications in either GRCh37 or GRCh38 result in reference-specific, missed variants for short- and long-read technologies in medically important genes including CBS, CRYAA, and KCNE1. Our proposed solution improves variant recall in these genes from 8% to 100%. This benchmark will significantly improve the comprehensive characterization of these medically relevant genes and guide new method development.

genomics↗

precisionFDA Truth Challenge V2: Calling variants from short- and long-reads in difficult-to-map regions

The precisionFDA Truth Challenge V2 aimed to assess the state-of-the-art of variant calling in difficult-to-map regions and the Major Histocompatibility Complex (MHC). Starting with FASTQ files, 20 challenge participants applied their variant calling pipelines and submitted 64 variant callsets for one or more sequencing technologies (~35X Illumina, ~35X PacBio HiFi, and ~50X Oxford Nanopore Technologies). Submissions were evaluated following best practices for benchmarking small variants with the new GIAB benchmark sets and genome stratifications. Challenge submissions included a number of innovative methods for all three technologies, with graph-based and machine-learning methods scoring best for short-read and long-read datasets, respectively. New methods out-performed the 2016 Truth Challenge winners, and new machine-learning approaches combining multiple sequencing technologies performed particularly well. Recent developments in sequencing and variant calling have enabled benchmarking variants in challenging genomic regions, paving the way for the identification of previously unknown clinically relevant variants.

bioinformatics↗