bioRxiv ScienceSearch

Biology subjects

Schatz, M.

Publications and source records attributed to Schatz, M..

5 recordsLinked to original sources

Skyhawk: An Artificial Neural Network-based discriminator for reviewing clinically significant genomic variants

MotivationMany rare diseases and cancers are fundamentally diseases of the genome. In the past several years, genome sequencing has become one of the most important tools in clinical practice for rare disease diagnosis and targeted cancer therapy. However, variant interpretation remains the bottleneck as is not yet automated and may take a specialist several hours of work per patient. On average, one-fifth of this time is spent on visually confirming the authenticity of the candidate variants. ResultsWe developed Skyhawk, an artificial neural network-based discriminator that mimics the process of expert review on clinically significant genomics variants. Skyhawk runs in less than one minute to review ten thousand variants, and about 30 minutes to review all variants in a typical whole-genome sequencing sample. Among the false positive singletons identified by GATK HaplotypeCaller, UnifiedGenotyper and 16GT in the HG005 GIAB sample, 79.7% were rejected by Skyhawk. Worked on the Variants with Unknown Significance (VUS), Skyhawk marked most of the false positive variants for manual review and most of the true positive variants no need for review. AvailabilitySkyhawk is easy to use and freely available at https://github.com/aquaskyline/Skyhawk

bioinformatics

Clairvoyante: a multi-task convolutional deep neural network for variant calling in Single Molecule Sequencing

The accurate identification of DNA sequence variants is an important, but challenging task in genomics. It is particularly difficult for single molecule sequencing, which has a per-nucleotide error rate of ~5%-15%. Meeting this demand, we developed Clairvoyante, a multi-task five-layer convolutional neural network model for predicting variant type (SNP or indel), zygosity, alternative allele and indel length from aligned reads. For the well-characterized NA12878 human sample, Clairvoyante achieved 99.73%, 97.68% and 95.36% precision on known variants, and 98.65%, 92.57%, 87.26% F1-score for whole-genome analysis, using Illumina, PacBio, and Oxford Nanopore data, respectively. Training on a second human sample shows Clairvoyante is sample agnostic and finds variants in less than two hours on a standard server. Furthermore, we identified 3,135 variants that are missed using Illumina but supported independently by both PacBio and Oxford Nanopore reads. Clairvoyante is available open-source (https://github.com/aquaskyline/Clairvoyante), with modules to train, utilize and visualize the model.

bioinformatics

REASSESSING THE REVOLUTIONS RESOLUTIONS

We are currently facing an avalanche of cryo-EM (cryogenic Electron Microscopy) publications presenting beautiful structures at resolution levels of ~3[A]: a true \"resolution revolution\" [Kuhlbrandt, Science 343(2014)1443-1444]. Impressive as these results may be, a fundamental statistical error has persisted in the literature that affects the numerical resolution values for practically all published structures. The error goes back to a misinterpretation of basic statistics and pervades virtually all popular cryo-EM quality metrics. The resolution in cryo-EM is typically assessed by the Fourier Shell Correlation \"FSC\" [Harauz & van Heel: Optik 73(1986)146-156] using a fixed threshold value of 0.143 (\"FSC0.143\") [Rosenthal, Henderson, J. Mol. Biol. 333(2003)721-745]. Using a simple model experiment we illustrate why this fixed threshold is flawed and we pinpoint the source of the resolution confusion. When two vectors are uncorrelated the expectation value of their inner-product is zero. That, however, does not imply that each individual inner-product of the vectors is zero (the vectors are not orthogonal). This error was introduced to electron microscopy in [Frank & Al-Ali, Nature 256(1975)376-379] and has since proliferated into virtually all quality and resolution-related metrics in EM. One criterion not affected by this error is the information-based [1/2]-bit FSC threshold [van Heel & Schatz: J. Struct. Biol. 151(2005)250-262].

biochemistry

Accurate detection of complex structural variations using single molecule sequencing

Structural variations (SVs) are the largest source of genetic variation, but remain poorly understood because of limited genomics technology. Single molecule long read sequencing from Pacific Biosciences and Oxford Nanopore has the potential to dramatically advance the field, although their high error rates challenge existing methods. Addressing this need, we introduce open-source methods for long read alignment (NGMLR, https://github.com/philres/ngmlr) and SV identification (Sniffles, https://github.com/fritzsedlazeck/Sniffles) that enable unprecedented SV sensitivity and precision, including within repeat-rich regions and of complex nested events that can have significant impact on human disorders. Examining several datasets, including healthy and cancerous human genomes, we discover thousands of novel variants using long reads and categorize systematic errors in short-read approaches. NGMLR and Sniffles are further able to automatically filter false events and operate on low amounts of coverage to address the cost factor that has hindered the application of long reads in clinical and research settings.

bioinformatics

Proper experimental design requires randomization/balancing of molecular ecology experiments

Properly designed (randomized and/or balanced) experiments are standard in ecological research. Molecular methods are increasingly used in ecology, but studies generally do not report the detailed design of sample processing in the laboratory. This may strongly influence the interpretability of results if the laboratory procedures do not account for the confounding effects of unexpected laboratory events. We demonstrate this with a simple experiment where unexpected differences in laboratory processing of samples would have biased results if randomization in DNA extraction and PCR steps do not provide safeguards. We emphasize the need for proper experimental design and reporting of the laboratory phase of molecular ecology research to ensure the reliability and interpretability of results.

ecology