bioRxiv ScienceSearch

Biology subjects

Kaplanis, J.

Publications and source records attributed to Kaplanis, J..

4 recordsLinked to original sources

Mutational origins and pathogenic consequences of multinucleotide mutations in 6,688 trios with developmental disorders

De novo mutations (DNMs) in protein-coding genes are a well-established cause of developmental disorders (DD). However, known DD-associated genes only account for a minority of the observed excess of such DNMs. To identify novel DD-associated genes, we integrated healthcare and research exome sequences on 31,058 DD parent-offspring trios, and developed a simulation-based statistical test to identify gene-specific enrichments of DNMs. We identified 299 significantly DD-associated genes, including 49 not previously robustly associated with DDs. Despite detecting more DD-associated genes than in any previous study, much of the excess of DNMs of protein-coding genes remains unaccounted for. Modelling suggests that over 500 novel DD-associated genes await discovery, many of which are likely to be less penetrant than the currently known genes. Research access to clinical diagnostic datasets will be critical for completing the map of dominant DDs.

genomics

Autosomal recessive coding variants explain only a small proportion of undiagnosed developmental disorders in the British Isles

Large exome-sequencing datasets offer an unprecedented opportunity to understand the genetic architecture of rare diseases, informing clinical genetics counseling and optimal study designs for disease gene identification. We analyzed 7,448 exome-sequenced families from the Deciphering Developmental Disorders study, and, for the first time, estimated the causal contribution of recessive coding variation exome-wide. We found that the proportion of cases attributable to recessive coding variants is surprisingly low in patients of European ancestry, at only 3.6%, versus 50% of cases explained by de novo coding mutations. Surprisingly, we found that, even in European probands with affected siblings, recessive coding variants are only likely to explain ~12% of cases. In contrast, they account for 31% of probands with Pakistani ancestry due to elevated autozygosity. We tested every gene for an excess of damaging homozygous or compound heterozygous genotypes and found three genes that passed stringent Bonferroni correction: EIF3F, KDM5B, and THOC6. EIF3F is a novel disease gene, and KDM5B has previously been reported as a dominant disease gene. KDM5B appears to follow a complex mode of inheritance, in which heterozygous loss-of-function variants (LoFs) show incomplete penetrance and biallelic LoFs are fully penetrant. Our results suggest that a large proportion of undiagnosed developmental disorders remain to be explained by other factors, such as noncoding variants and polygenic risk.

genetics

Quantitative analysis of population-scale family trees using millions of relatives

Family trees have vast applications in multiple fields from genetics to anthropology and economics. However, the collection of extended family trees is tedious and usually relies on resources with limited geographical scope and complex data usage restrictions. Here, we collected 86 million profiles from publicly-available online data from genealogy enthusiasts. After extensive cleaning and validation, we obtained population-scale family trees, including a single pedigree of 13 million individuals. We leveraged the data to partition the genetic architecture of longevity by inspecting millions of relative pairs and to provide insights to population genetics theories on the dispersion of families. We also report a simple digital procedure to overlay other datasets with our resource in order to empower studies with population-scale genealogical data.\n\nOne Sentence SummaryUsing massive crowd-sourced genealogy data, we created a population-scale family tree resource for scientific studies.

genomics

Striking differences in patterns of germline mutation between mice and humans

Recent whole genome sequencing (WGS) studies have estimated that the human germline mutation rate per basepair per generation ([~]1.2-10-8) 1,2 is substantially higher than in mice (3.5-5.4-10-9) 3,4, which has been attributed to more efficient purifying selection due to larger effective population sizes in mice compared to humans.5,6,7. In humans, most germline mutations are paternal in origin and the numbers of mutations per offspring increase markedly with paternal age 2,8,9 and more weakly with maternal age 10. Germline mutations can arise at any stage of the cellular lineage from zygote to gamete, resulting in mutations being represented in different proportion and types of cells, with the earliest embryonic mutations being mosaic in both somatic and germline cells. Here we use WGS of multi-sibling mouse and human pedigrees to show striking differences in germline mutation rate and spectra between the two species, including a dramatic reduction in mutation rate in human spermatogonial stem cell (SSC) divisions, which we hypothesise was driven by selection. The differences we observed between mice and humans result from both biological differences within the same stage of embryogenesis or gametogenesis and species-specific differences in cellular genealogies of the germline.

genomics