bioRxiv Science⌕ Search

Biology subjects

Mehaffey, M.

Publications and source records attributed to Mehaffey, M..

2 recordsLinked to original sources

Genome-wide maps of highly-similar intrachromosomal repeats that mediate ectopic recombination in three human genome assemblies

Repeated sequences spread throughout the genome play important roles in shaping the structure of chromosomes and facilitating the generation of new genomic variation. Through a variety of mechanisms, repeats are involved in generating structural rearrangements such as deletions, duplications, inversions, and translocations, which can have the potential to impact human health. Despite their significance, repetitive regions including tandem repeats, transposable elements, segmental duplications, and low-copy repeats remain a challenge to characterize due to technological limitations inherent to many sequencing methodologies. We performed genome-wide analyses and comparisons of direct and inverted repeated sequences in the latest available human genome reference assemblies including GRCh37 and GRCh38 and the most recent telomere-to-telomere alternate assembly (T2T-CHM13). Overall, the composition and distribution of direct and inverted repeats identified remains similar among the three assemblies but we observed an increase in the number of repeated sequences detected in the T2T-CHM13 assembly versus the reference assemblies. As expected, there is an enrichment of repetitive regions in the short arms of acrocentric chromosomes, which had been previously unresolved in the human genome reference assemblies. We cross-referenced the identified repeats with protein-coding genes across the genome to identify those at risk for being involved in genomic disorders. We observed that certain gene categories, such as olfactory receptors and immune response genes, are enriched among those impacted by repeated sequences likely contributing to human diversity and adaptation. Through this analysis, we have produced a catalogue of direct and inversely oriented repeated sequences across the currently three most widely used human genome assemblies. Bioinformatic analyses of these repeats and their contribution to genome architecture can reveal regions that are most susceptible to genomic instability. Understanding how the architectural genomic features of repeat pairs such as their homology, size and distance can lead to complex genomic rearrangement formation can provide further insights into the molecular mechanisms leading to genomic disorders and genome evolution. Author summaryThis study focused on the characterization of intrachromosomal repeated sequences in the human genome that can play important roles in shaping chromosome structure and generating new genomic variation in three human genome assemblies. We observed an increase in the number of repeated sequence pairs detected in the most recent telomere-to-telomere alternate assembly (T2T-CHM13) compared to the reference assemblies (GRCh37 and GRCh38). We observed an enrichment of repeats in the T2T-CHM13 acrocentric chromosomes, which had been previously unresolved. Importantly, our study provides a catalogue of direct and inverted repeated sequences across three commonly used human genome assemblies, which can aid in the understanding of genomic architecture instability, evolution, and disorders. Our analyses provide insights into repetitive regions in the human genome that may contribute to complex genomic rearrangements

genomics↗

Targeted analysis of dyslexia-associated regions on chromosomes 6, 12 and 15 in large multigenerational cohorts

Dyslexia is a common specific learning disability with a strong genetic basis that affects word reading and spelling. An increasing list of loci and genes have been implicated, but analyses to-date investigated only limited genomic variation within each locus with no confirmed pathogenic variants. In a collection of >2000 participants in families enrolled at three independent sites, we performed targeted capture and comprehensive sequencing of all exons and some regulatory elements of five candidate dyslexia risk genes (DNAAF4, CYP19A1, DCDC2, KIAA0319 and GRIN2B) for which prior evidence of association exists from more than one sample. For each of six dyslexia-related phenotypes we used both individual-single nucleotide polymorphism (SNP) and aggregate testing of multiple SNPs to evaluate evidence for association. We detected no promoter alterations and few potentially deleterious variants in the coding exons, none of which showed evidence of association with any phenotype. All genes except DNAAF4 provided evidence of association, corrected for the number of genes, for multiple non-coding variants with one or more phenotypes. Results for a variant in the downstream region of CYP19A1 and a haplotype in DCDC2 yielded particularly strong statistical significance for association. This haplotype and another in DCDC2 affected performance of real word reading in opposite directions. In KIAA0319, two missense variants annotated as tolerated/benign associated with poor performance on spelling. Ten non-coding SNPs likely affect transcription factor binding. Findings were similar regardless of whether phenotypes were adjusted for verbal IQ. Our findings from this large-scale sequencing study complement those from genome-wide association studies (GWAS), argue strongly against the causative involvement of large-effect coding variants in these five candidate genes, support an oligogenic etiology, and suggest a role of transcriptional regulation. Author SummaryFamily studies show that genes play a role in dyslexia and a small number of genomic regions have been implicated to date. However, it has proven difficult to identify the specific genetic variants in those regions that affect reading ability by using indirect measures of association with evenly spaced polymorphisms chosen without regard to likely function. Here, we use recent advances in DNA sequencing to examine more comprehensively the role of genetic variants in five previously nominated candidate dyslexia risk genes on several dyslexia-related traits. Our analysis of more than 2000 participants in families with dyslexia provides strong evidence for a contribution to dyslexia risk for the non-protein coding genetic variant rs9930506 in the CYP19A1 gene on chromosome 15 and excludes the DNAAF4 gene on the same chromosome. We identified other putative causal variants in genes DCDC2 and KIAA0319 on chromosome 6 and GRIN2B on chromosome 12. Further studies of these DNA variants, all of which were non-coding, may point to new biological pathways that affect susceptibility to dyslexia. These findings are important because they implicate regulatory variation in this complex trait that affects ability of individuals to effectively participate in our increasingly informatic world.

genetics↗