bioRxiv ScienceSearch

Biology subjects

Wenger, A. M.

Publications and source records attributed to Wenger, A. M..

5 recordsLinked to original sources

Phrank measures phenotype sets similarity to greatly improve Mendelian diagnostic disease prioritization

PurposeExome sequencing and diagnosis is beginning to spread across the medical establishment. The most time-consuming part of genome based diagnosis is the manual step of matching the potentially long list of patient candidate genes to patient phenotypes to identify the causative disease.\n\nMethodsWe introduce Phrank (for phenotype ranking), an information-theory inspired method that utilizes a Bayesian Network to prioritize candidate diseases or genes, as a stand-alone module that can be run with any underlying knowledgebase and any variant filtering scheme.\n\nResultsPhrank outperforms existing methods at ranking the causative disease or gene when applied to 169 real patient exomes with Mendelian diagnoses. Phranks greatest improvement is in disease space, where across all 169 patients it ranks only 3 diseases on average ahead of the true diagnosis, whereas Phenomizer ranks 32 diseases ahead of the causal one.\n\nConclusionUsing Phrank to rank all patient candidate genes or diseases, as they start working through a new case, will save the busy clinician much time in deriving a genetic diagnosis.

genomics

Independent erosion of conserved transcription factor binding sites points to shared hindlimb, vision, and scrotum loss in different mammals

Genetic variation in cis-regulatory elements is thought to be a major driving force in morphological and physiological change. However, identifying transcription factor binding events which code for complex traits remains a challenge, motivating novel means of detecting putatively important binding events. Using a curated set of 1,154 high-quality transcription factor motifs, we demonstrate that independently eroded binding sites are enriched for independently lost traits in three distinct pairs of placental mammals. We show that these independently eroded events pinpoint the loss of hindlimbs in dolphin and manatee, degradation of vision in naked mole-rat and star-nosed mole, and the loss of scrotum in white rhinoceros and Weddell seal. Our study exhibits a novel methodology to detect cis-regulatory mutations which help explain a portion of the molecular mechanism underlying complex trait formation and loss.\n\nAuthor SummaryEvolution has produced an astounding variety of species with incredibly diverse phenotypes. A central question in evolutionary developmental biology is how (and which) DNA evolves to encode all of these different traits. A prevailing hypothesis is that changes in regulatory DNA, short stretches of DNA which control the expression of protein-coding genes, drive important differences in trait formation between species. The basic building block of regulatory DNA is thought to be transcription factor binding sites, shortl genomic sequences which attract proteins whose central role is to control the rate of transcription. In this study, we asked whether the independent erosion of otherwise highly conserved transcription factor binding sites points to a trait shared between species which have undergone similar adaptations. We show that our method is able to point to the loss of hindlimbs in dolphin and manatee, poor vision in naked mole-rat and star-nosed mole, and loss of scrotum in Weddell seal and white rhinoceros. Overall, our study exhibits a means of detecting evolutionarily important genomic regions which help explain a portion of complex trait loss and retention.

evolutionary biology

Multi-platform discovery of haplotype-resolved structural variation in human genomes

The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, and strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three human parent-child trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs ([&ge;]50 bp) per human genome. We also discover 156 inversions per genome--most of which previously escaped detection. Fifty-eight of the inversions we discovered intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The method and the dataset serve as a gold standard for the scientific community and we make specific recommendations for maximizing structural variation sensitivity for future large-scale genome sequencing studies.

genomics

AMELIE accelerates Mendelian patient diagnosis directly from the primary literature

The diagnosis of Mendelian disorders requires labor-intensive literature research. Our software system AMELIE (Automatic Mendelian Literature Evaluation) greatly automates this process. AMELIE parses hundreds of thousands of full text articles to find an underlying diagnosis to explain a patients phenotypes given the patients exome. AMELIE prioritizes patient candidate genes for their likelihood of causing the patients phenotypes. Diagnosis of singleton patients (without relatives exomes) is the most time-consuming scenario. AMELIEs gene ranking method was tested on 215 singleton Mendelian patients with a clinical diagnosis. AMELIE ranked the causal gene among the top 2 in the majority (63%) of cases. Examining AMELIEs top 10 genes, amounting to 8% of 124 candidate genes with rare functional variants per patient, results in diagnosis for 95% of cases. Strikingly, training only on gene pathogenicity knowledge from 2011 leads to identical performance compared to training on current data. An accompanying analysis web portal has launched at AMELIE.stanford.edu.

genetics

Long-read whole genome sequencing identifies causal structural variation in a Mendelian disease

Current clinical genomics assays primarily utilize short-read sequencing (SRS), which offers high throughput, high base accuracy, and low cost per base. SRS has, however, limited ability to evaluate tandem repeats, regions with high [GC] or [AT] content, highly polymorphic regions, highly paralogous regions, and large-scale structural variants. Long-read sequencing (LRS) has complementary strengths and offers a means to discover overlooked genetic variation in patients undiagnosed by SRS. To evaluate LRS, we selected a patient who presented with multiple neoplasia and cardiac myxomata suggestive of Carney complex for whom targeted clinical gene testing and whole genome SRS were negative. Low coverage whole genome LRS was performed on the PacBio Sequel system and structural variants were called, yielding 6,971 deletions and 6,821 insertions > 50bp. Filtering for variants that are absent in an unrelated control and that overlap a coding exon of a disease gene identified three deletions and three insertions. One of these, a heterozygous 2,184 bp deletion, overlaps the first coding exon of PRKAR1A, which is implicated in autosomal dominant Carney complex. This variant was confirmed by Sanger sequencing and was classified as pathogenic using standard criteria for the interpretation of sequence variants. This first successful application of whole genome LRS to identify a pathogenic variant suggests that LRS has significant potential to identify disease-causing structural variation. We recommend larger studies to evaluate the diagnostic yield of LRS, and the development of a comprehensive catalog of common human structural variation to support future studies.

genomics