bioRxiv ScienceSearch

bioRxiv · 10.1101/016618

Two variance component model improves genetic prediction in family data sets

Abstract

Genetic prediction based on either identity by state (IBS) sharing or pedigree information has been investigated extensively using Best Linear Unbiased Prediction (BLUP) methods. Such methods were pioneered in the plant and animal breeding literature and have since been applied to predict human traits with the aim of eventual clinical utility. However, methods to combine IBS sharing and pedigree information for genetic prediction in humans have not been explored. We introduce a two variance component model for genetic prediction: one component for IBS sharing and one for approximate pedigree structure, both estimated using genetic markers. In simulations using real genotypes from CARe and FHS family cohorts, we demonstrate that the two variance component model achieves gains in prediction r2 over standard BLUP at current sample sizes, and we project based on simulations that these gains will continue to hold at larger sample sizes. Accordingly, in analyses of four quantitative phenotypes from CARe and two quantitative phenotypes from FHS, the two variance component model significantly improves prediction r2 in each case, with up to a 20% relative improvement. We also find that standard mixed model association tests can produce inflated test statistics in data sets with related individuals, whereas the two variance component model corrects for inflation.\n\nAuthor SummaryGenetic prediction has been well-studied in plant and animal breeding and has generated considerable recent interest in human genetics, both in family data sets and in population cohorts. Many prediction studies are based on the widely used Best Linear Unbiased Prediction (BLUP) approach, which performs a mixed model analysis using a genetic relationship matrix that is either estimated from genotype data--thus measuring identity-by-state (IBS) sharing--or obtained from family pedigree information. We show here that a substantial improvement in prediction accuracy in family data sets can be obtained by jointly modeling both IBS sharing and approximate pedigree structure, both estimated using genetic markers, using separate variance components within a two variance component mixed model. We demonstrate the performance of this model in simulations and real data sets. We also show that previous mixed model association methods suffer from inflated test statistics in family data sets due to their failure to account for the different heritability parameters corresponding to IBS sharing vs. pedigree relatedness. Our two variance component model provides a solution to this problem without compromising statistical power.

Explore related subjects

Keep this discovery

BibTeXRIS

George Tucker, Po-Ru Loh, Iona M MacLeod, Ben J Hayes, Michael E Goddard, Bonnie Berger, Alkes L Price. 2015-03-17. Two variance component model improves genetic prediction in family data sets. https://doi.org/10.1101/016618

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Rapid evolution of primate type 2 immune response factors linked to asthma susceptibility

Host immunity pathways evolve rapidly in response to antagonism by pathogens. Microbial infections can also trigger excessive inflammation that contributes to diverse autoimmune disorders including asthma, lupus, diabetes, and arthritis. Definitive links between immune system evolution and human autoimmune disease remain unclear. Here we provide evidence that several components of the type 2 immune response pathway have been subject to recurrent positive selection in the primate lineage. Notably, rapid evolution of the central immune regulator IL13 corresponds to a polymorphism linked to asthma susceptibility in humans. We also find evidence of accelerated amino acid substitutions as well as repeated gene gain and loss events among eosinophil granule proteins, which act as toxic antimicrobial effectors that promote asthma pathology by damaging airway tissues. These results support the hypothesis that evolutionary conflicts with pathogens promote tradeoffs for increasingly robust immune responses during animal evolution. Our findings are also consistent with the view that natural selection has contributed to the spread of autoimmune disease alleles in humans.

Genetics

Single Cell Expression Data Reveal Human Genes that Escape X-Chromosome Inactivation

Sex chromosomes pose an inherent genetic imbalance between genders. In mammals, one of the females X-chromosomes undergoes inactivation (Xi). Indirect measurements estimate that about 20% of Xi genes completely or partially escape inactivation. The identity of these escapee genes and their propensity to escape inactivation remain unsolved. A direct method for identifying escapees was applied by quantifying differential allelic expression from single cells. RNA-Seq fragments were assigned to informative SNPs which were labeled by the appropriate parental haplotype. This method was applied for measuring allelic specific expression from Chromosome-X (ChrX) and an autosomal chromosome as a control. We applied the protocol for measuring biallelic expression from ChrX to 104 primary fibroblasts. Out of 215 genes that were considered, only 13 genes (6%) were associated with biallelic expression. The sensitivity of escapees' identification was increased by combining SNP mapping for parental diploid genomes together with RNA-Seq from clonal single cells (25 lymphoblasts). Using complementary protocols, referred to as strict and relaxed, we confidently identified 25 and 31escapee genes, respectively. When pooled versions of 30 and 100 cells were used, <50% of these genes were revealed. We assessed the generality of our protocols in view of an escapee catalog compiled from indirect methods. The overlap between the escapee catalog and the genes list from this study is statistically significant (P-value of E-07). We conclude that single cells expression data are instrumental for studying X-inactivation with an improved sensitivity. Finally, our results support the emerging notion of the non-deterministic nature of genes that escape X-chromosome inactivation.

Genetics

Frequency of mosaicism points towards mutation-prone early cleavage cell divisions.

It has recently become possible to directly estimate the germ-line de novo mutation (dnm) rate by sequencing the whole genome of father-mother-offspring trios, and this has been conducted in human1-5, chimpanzee6, mice7, birds8 and fish9. In these studies dnms are typically defined as variants that are heterozygous in the offspring while being absent in both parents. They are assumed to have occurred in the germ-line of one of the parents and to have been transmitted to the offspring via the sperm cell or oocyte. This definition assumes that detectable mosaicism in the parent in which the mutation occurred is negligible. However, instances of detectable mosaicism or premeiotic clusters are well documented in humans and other organisms, including ruminants10-12. We herein take advantage of cattle pedigrees to show that as much as [~]30% to [~]50% of dnms present in a gamete may occur during the early cleavage cell divisions in males and females, respectively, resulting in frequent detectable mosaicism and a high rate of sharing of multiple dnms between siblings. This should be taken into account to accurately estimate the mutation rate in cattle and other species.

Genetics