bioRxiv ScienceSearch

Biology subjects

Gymrek, M.

Publications and source records attributed to Gymrek, M..

6 recordsLinked to original sources

Profiling the genome-wide landscape of tandem repeat expansions

Tandem Repeat (TR) expansions have been implicated in dozens of genetic diseases, including Huntingtons Disease, Fragile X Syndrome, and hereditary ataxias. Furthermore, TRs have recently been implicated in a range of complex traits, including gene expression and cancer risk. While the human genome harbors hundreds of thousands of TRs, analysis of TR expansions has been mainly limited to known pathogenic loci. A major challenge is that expanded repeats are beyond the read length of most next-generation sequencing (NGS) datasets. We present GangSTR, a novel algorithm for genome-wide profiling of both normal and expanded TRs. GangSTR extracts information from paired-end reads into a unified model to estimate maximum likelihood TR lengths. We validated GangSTR on real and simulated TR expansions and show that GangSTR outperforms alternative methods. We applied GangSTR to more than 150 individuals to profile the landscape of TR expansions in a healthy population and validated novel expansions using orthogonal technologies. Our analysis revealed that each individual harbors dozens of TR alleles longer than standard read lengths and identified hundreds of potentially mis-annotated TRs in the reference genome. GangSTR is packaged as a standalone tool that will likely enable discovery of novel pathogenic variants not currently accessible from NGS.

genomics

A reference haplotype panel for genome-wide imputation of short tandem repeats

Short tandem repeats (STRs) are involved in dozens of Mendelian disorders and have been implicated in a variety of complex traits. However, existing technologies focusing on single nucleotide polymorphisms (SNPs) have not allowed for systematic STR association studies. Here, we leverage next-generation sequencing data from 479 families to create a SNP+STR reference haplotype panel for genome-wide imputation of STRs into SNP data. Imputation achieved an average of 97% concordance between genotyped and imputed STR genotypes in an external dataset compared to 63% expected under a random model. Performance varied widely across STRs, with near perfect concordance at bi-allelic STRs vs. 70% at highly polymorphic forensics markers. We demonstrate that imputation increases power over individual SNPs to detect STR associations using simulated phenotypes and gene expression data. This resource will enable the first large-scale STR association studies using existing SNP datasets, and will likely yield new insights into complex traits.

genomics

Targeted Genotyping of Variable Number Tandem Repeats with adVNTR

Whole Genome Sequencing is increasingly used to identify Mendelian variants in clinical pipelines. These pipelines focus on single nucleotide variants (SNVs) and also structural variants, while ignoring more complex repeat sequence variants. We consider the problem of genotyping Variable Number Tandem Repeats (VNTRs), composed of inexact tandem duplications of short (6-100bp) repeating units. VNTRs span 3% of the human genome, are frequently present in coding regions, and have been implicated in multiple Mendelian disorders. While existing tools recognize VNTR carrying sequence, genotyping VNTRs (determining repeat unit count and sequence variation) from whole genome sequenced reads remains challenging. We describe a method, adVNTR, that uses Hidden Markov Models to model each VNTR, count repeat units, and detect sequence variation. adVNTR models can be developed for short-read (Illumina) and single molecule (PacBio) whole genome and exome sequencing, and show good results on multiple simulated and real data sets. adVNTR is available at https://github.com/mehrdadbakhtiari/adVNTR

bioinformatics

Quantification of autism recurrence risk by direct assessment of paternal sperm mosaicism

De novo genetic mutations represent a major contributor to pediatric disease, including autism spectrum disorders (ASD), congenital heart disease, and muscular dystrophies1,2, but there are currently no methods to prevent or predict them. These mutations are classically thought to occur either at low levels in progenitor cells or at the time of fertilization1,3 and are often assigned a low risk of recurrence in siblings4,5. Here, we directly assess the presence of de novo mutations in paternal sperm and discover abundant, germline-restricted mosaicism. From a cohort of ASD cases, employing single molecule genotyping, we found that four out of 14 fathers were germline mosaic for a putatively causative mutation transmitted to the affected child. Three of these were enriched or exclusively present in sperm at high allelic fractions (AF; 7-15%); and one was recurrently transmitted to two additional affected children, representing clinically actionable information. Germline mosaicism was further assessed by deep (>90x) whole genome sequencing of four paternal sperm samples, which detected 12/355 transmitted de novo single nucleotide variants that were mosaic above 2% AF, and more than two dozen additional, non-transmitted mosaic variants in paternal sperm. Our results demonstrate that germline mosaicism is an underestimated phenomenon, which has important implications for clinical practice and in understanding the basis of human disease. Genetic analysis of sperm can assess individualized recurrence risk following the birth of a child with a de novo disease, as well as the risk in any male planning to have children.

genetics

Quantitative analysis of population-scale family trees using millions of relatives

Family trees have vast applications in multiple fields from genetics to anthropology and economics. However, the collection of extended family trees is tedious and usually relies on resources with limited geographical scope and complex data usage restrictions. Here, we collected 86 million profiles from publicly-available online data from genealogy enthusiasts. After extensive cleaning and validation, we obtained population-scale family trees, including a single pedigree of 13 million individuals. We leveraged the data to partition the genetic architecture of longevity by inspecting millions of relative pairs and to provide insights to population genetics theories on the dispersion of families. We also report a simple digital procedure to overlay other datasets with our resource in order to empower studies with population-scale genealogical data.\n\nOne Sentence SummaryUsing massive crowd-sourced genealogy data, we created a population-scale family tree resource for scientific studies.

genomics

A framework to interpret short tandem repeat variation in humans

Identifying regions of the genome that are depleted of mutations can reveal potentially deleterious variants. Short tandem repeats (STRs), also known as microsatellites, are among the largest contributors of de novo mutations in humans and are implicated in a variety of human disorders. However, because of the challenges STRs pose to bioinformatics tools, per-locus studies of STR mutations have been limited to highly ascertained panels of several dozen loci. Here, we harnessed bioinformatics tools and a novel analytical framework to estimate mutation parameters for each STR in the human genome by correlating STR genotypes with local sequence heterozygosity. We applied our method to obtain robust estimates of the impact of local sequence features on mutation parameters and used this to create a framework for measuring constraint at STRs by comparing observed vs. expected mutation rates. Constraint scores identified known pathogenic variants with early onset effects. Our constraint metrics will provide a valuable tool for prioritizing pathogenic STRs in medical genetics studies.

genetics