bioRxiv ScienceSearch

bioRxiv · 10.1101/025379

Construction of relatedness matrices using genotyping-by-sequencing data

Abstract

BackgroundGenotyping-by-sequencing (GBS) is becoming an attractive alternative to array-based methods for genotyping individuals for a large number of single nucleotide polymorphisms (SNPs). Costs can be lowered by reducing the mean sequencing depth, but this results in genotype calls of lower quality. A common analysis strategy is to filter SNPs to just those with sufficient depth, thereby greatly reducing the number of SNPs available. We investigate methods for estimating relatedness using GBS data, including results of low depth, using theoretical calculation, simulation and application to a real data set.\n\nResultsWe show that unbiased estimates of relatedness can be obtained by using only those SNPs with genotype calls in both individuals. The expected value of this estimator is independent of the SNP depth in each individual, under a model of genotype calling that includes the special case of the two alleles being read at random. In contrast, the estimator of self-relatedness does depend on the SNP depth, and we provide a modification to provide unbiased estimates of self-relatedness. We refer to these methods of estimation as kinship using GBS with depth adjustment (KGD). The estimators can be calculated using matrix methods, which allow efficient computation. Simulation results were consistent with the methods being unbiased, and suggest that the optimal sequencing depth is around 2-4 for relatedness between individuals and 5-10 for self-relatedness. Application to a real data set revealed that some SNP filtering may still be necessary, for the exclusion of SNPs which did not behave in a Mendelian fashion. A simple graphical method (a fin plot) is given to illustrate this issue and to guide filtering parameters.\n\nConclusionWe provide a method which gives unbiased estimates of relatedness, based on SNPs assayed by GBS, which accounts for the depth (including zero depth) of the genotype calls. This allows GBS to be applied at read depths which can be chosen to optimise the information obtained. SNPs with excess heterozygosity, often due to (partial) polyploidy or other duplications can be filtered based on a simple graphical method.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Ken G Dodds, John C McEwan, Rudiger Brauning, Rayna M Anderson, Tracey C van Stijn, Theodor Kristjánsson, Shannon M Clarke. 2015-08-24. Construction of relatedness matrices using genotyping-by-sequencing data. https://doi.org/10.1101/025379

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Rapid evolution of primate type 2 immune response factors linked to asthma susceptibility

Host immunity pathways evolve rapidly in response to antagonism by pathogens. Microbial infections can also trigger excessive inflammation that contributes to diverse autoimmune disorders including asthma, lupus, diabetes, and arthritis. Definitive links between immune system evolution and human autoimmune disease remain unclear. Here we provide evidence that several components of the type 2 immune response pathway have been subject to recurrent positive selection in the primate lineage. Notably, rapid evolution of the central immune regulator IL13 corresponds to a polymorphism linked to asthma susceptibility in humans. We also find evidence of accelerated amino acid substitutions as well as repeated gene gain and loss events among eosinophil granule proteins, which act as toxic antimicrobial effectors that promote asthma pathology by damaging airway tissues. These results support the hypothesis that evolutionary conflicts with pathogens promote tradeoffs for increasingly robust immune responses during animal evolution. Our findings are also consistent with the view that natural selection has contributed to the spread of autoimmune disease alleles in humans.

Genetics

Single Cell Expression Data Reveal Human Genes that Escape X-Chromosome Inactivation

Sex chromosomes pose an inherent genetic imbalance between genders. In mammals, one of the females X-chromosomes undergoes inactivation (Xi). Indirect measurements estimate that about 20% of Xi genes completely or partially escape inactivation. The identity of these escapee genes and their propensity to escape inactivation remain unsolved. A direct method for identifying escapees was applied by quantifying differential allelic expression from single cells. RNA-Seq fragments were assigned to informative SNPs which were labeled by the appropriate parental haplotype. This method was applied for measuring allelic specific expression from Chromosome-X (ChrX) and an autosomal chromosome as a control. We applied the protocol for measuring biallelic expression from ChrX to 104 primary fibroblasts. Out of 215 genes that were considered, only 13 genes (6%) were associated with biallelic expression. The sensitivity of escapees' identification was increased by combining SNP mapping for parental diploid genomes together with RNA-Seq from clonal single cells (25 lymphoblasts). Using complementary protocols, referred to as strict and relaxed, we confidently identified 25 and 31escapee genes, respectively. When pooled versions of 30 and 100 cells were used, <50% of these genes were revealed. We assessed the generality of our protocols in view of an escapee catalog compiled from indirect methods. The overlap between the escapee catalog and the genes list from this study is statistically significant (P-value of E-07). We conclude that single cells expression data are instrumental for studying X-inactivation with an improved sensitivity. Finally, our results support the emerging notion of the non-deterministic nature of genes that escape X-chromosome inactivation.

Genetics

Frequency of mosaicism points towards mutation-prone early cleavage cell divisions.

It has recently become possible to directly estimate the germ-line de novo mutation (dnm) rate by sequencing the whole genome of father-mother-offspring trios, and this has been conducted in human1-5, chimpanzee6, mice7, birds8 and fish9. In these studies dnms are typically defined as variants that are heterozygous in the offspring while being absent in both parents. They are assumed to have occurred in the germ-line of one of the parents and to have been transmitted to the offspring via the sperm cell or oocyte. This definition assumes that detectable mosaicism in the parent in which the mutation occurred is negligible. However, instances of detectable mosaicism or premeiotic clusters are well documented in humans and other organisms, including ruminants10-12. We herein take advantage of cattle pedigrees to show that as much as [~]30% to [~]50% of dnms present in a gamete may occur during the early cleavage cell divisions in males and females, respectively, resulting in frequent detectable mosaicism and a high rate of sharing of multiple dnms between siblings. This should be taken into account to accurately estimate the mutation rate in cattle and other species.

Genetics