bioRxiv ScienceSearch

Biology subjects

Sarmady, M.

Publications and source records attributed to Sarmady, M..

3 recordsLinked to original sources

Genetic variant pathogenicity prediction trained using large-scale disease specific clinical sequencing datasets

Recent advances in DNA sequencing technologies have expanded our understanding of the molecular underpinnings for several genetic disorders, and increased the utilization of genomic tests by clinicians. Given the paucity of evidence to assess each variant, and the difficulty of experimentally evaluating a variants clinical significance, many of the thousand variants that can be generated by clinical tests are reported as variants of unknown clinical significance. However, the creation of population-scale variant databases can significantly improve clinical variant interpretation. Specifically, pathogenicity prediction for novel missense variants can now utilize features describing regional variant constraint. Constrained genomic regions are those that have an unusually low variant count in the general population. Several computational methods have been introduced to capture these regions and incorporate them into pathogenicity classifiers, but these methods have yet to be compared on an independent clinical variant dataset. Here we introduce one variant dataset derived from clinical sequencing panels, and use it to compare the ability of different genomic constraint metrics to determine missense variant pathogenicity. This dataset is compiled from 17,071 patients surveyed with clinical genomic sequencing for cardiomyopathy, epilepsy, or RASopathies. We further utilize this dataset to demonstrate the necessity of disease-specific classifiers, and to train PathoPredictor, a disease-specific ensemble classifier of pathogenicity based on regional constraint and variant level features. PathoPredictor achieves an average precision greater than 90% for variants from all 99 tested disease genes while approaching 100% accuracy for some genes. Accumulation of larger clinical variant datasets and their utilization to train existing pathogenicity metrics can significantly enhance their performance in a disease and gene-specific manner.

genomics

Rapid interpretation of clinical exomes using Phenoxome: a computational phenotype-driven approach

Clinical exome sequencing (CES) has become the preferred diagnostic platform for complex pediatric disorders with suspected monogenic etiologies, solving up to 20%-50% of cases depending on indication. Despite rapid advancements in CES analysis, the major challenge still resides in identifying the casual variants among the thousands of variants detected during CES testing, and thus establishing a molecular diagnosis. To improve the clinical exome diagnostic efficiency, we developed Phenoxome, a robust phenotype-driven model that adopts a network-based approach to facilitate automated variant prioritization and subsequent classification. Phenoxome dissects the phenotypic manifestation of a patient in conjunction with their genomic profile to filter and then prioritize putative pathogenic variants. To validate our method, we have compiled a clinical cohort of 105 positive patient samples (i.e. at least one reported pathogenic variant) that represent a wide range of genetic heterogeneity from The Childrens Hospital of Philadelphia. Our approach identifies the causative variants within the top 5, 10, or 25 candidates in more than 50%, 71%, or 88% of these patient samples respectively. Furthermore, we show that our method is optimized for clinical testing by yielding superior ranking of the pathogenic variants compared to current state-of-art methods. The web application of Phenoxome is available to the public at http://phenoxome.chop.edu/.

bioinformatics

ExomeSlicer: a resource for the development and validation of exome-based clinical panels

Exome-based panels (exome slices) are becoming the preferred diagnostic strategy in clinical laboratories, especially for genetically heterogeneous disorders. The advantages of this approach include enabling frequent updates to gene content without the need for re-designing, reflexing to exome analysis bioinformatically without requiring additional sequencing, and streamlining laboratory operation by using established exome kits and protocols. Despite their increasing use, there are currently no guidelines or appropriate resources to support their clinical implementation. Here, we highlight principles and important considerations for the clinical development and validation of exome-based panels, guided by clinical data from a diagnostic epilepsy panel using this approach. We also present a novel, publically accessible web-based resource, ExomeSlicer, and demonstrate its clinical utility in predicting gene-specific and exome-wide technically challenging regions that are not amenable to Next Generation Sequencing (NGS), and that might significantly lead to increased post hoc Sanger fill in burden. Using this tool, we also characterize > 2000 low complexity, GC-rich and/or high homology, regions across the exome that can be a source of false positive or false negative variant calls thus potentially leading to misdiagnoses in tested patients.

genomics