bioRxiv ScienceSearch

Biology subjects

Eran Halperin

Publications and source records attributed to Eran Halperin.

6 recordsLinked to original sources

Genome-wide methylation data mirror ancestry information

Genetic data are known to harbor information about human demographics, and genotyping data are commonly used for capturing ancestry information by leveraging genome-wide differences between populations. In contrast, it is not clear to what extent population structure is captured by whole-genome DNA methylation data. We demonstrate, using three large cohort 450K methylation array data sets, that ancestry information signal is mirrored in genome-wide DNA methylation data, and that it can be further isolated more effectively by leveraging the correlation structure of CpGs with cis-located SNPs. Based on these insights, we propose a method, EPISTRUCTURE, for the inference of ancestry from methylation data, without the need for genotype data. EPISTRUCTURE can be used to infer ancestry information of individuals based on their methylation data in the absence of corresponding genetic data. Although genetic data are often collected in epigenetic studies of large cohorts, these are typically not made publicly available, making the application of EPISTRUCTURE especially useful for anyone working on public data. Implementation of EPISTRUCTURE is available in GLINT, our recently released toolset for DNA methylation analysis at: http://glint-epigenetics.readthedocs.io.

Genetics

The Effects of Migration and Assortative Mating on Admixture Linkage Disequilibrium

1Statistical models in medical and population genetics typically assume that individuals assort randomly in a population. While this simplifies model complexity, it contradicts an increasing body of evidence of non-random mating in human populations. Specifically, it has been shown that assortative mating is significantly affected by genomic ancestry. In this work we examine the effects of ancestry-assortative mating on the linkage disequilibrium between local ancestry tracks of individuals in an admixed population. To accomplish this, we develop an extension to the Wright-Fisher model that allows for ancestry based assortative mating. We show that ancestry-assortment perturbs the distribution of local ancestry linkage disequilibrium (LAD) and the variance of ancestry in a population as a function of the number of generations since admixture. This assortment effect can induce errors in demographic inference of admixed populations when methods assume random mating. We derive closed form formulae for LAD under an assortative-mating model with and without migration. We observe that LAD depends on the correlation of global ancestry of couples in each generation, the migration rate of each of the ancestral populations, the initial proportions of ancestral populations, and the number of generations since admixture. We also present the first evidence of ancestry-assortment in African Americans and examine LAD in simulated and real admixed population data of African Americans. We find that demographic inference under the assumption of random mating significantly underestimates the number of generations since admixture, and that accounting for assortative mating using the patterns of LAD results in estimates that more closely agrees with the historical narrative.

Genetics

An Ancestry Based Approach for Detecting Interactions

IBackgroundEpistasis and gene-environment interactions are known to contribute significantly to variation of complex phenotypes in model organisms. However, their identification in human association studies remains challenging for myriad reasons. In the case of epistatic interactions, the large number of potential interacting sets of genes presents computational, multiple hypothesis correction, and other statistical power issues. In the case of gene-environment interactions, the lack of consistently measured environmental covariates in most disease studies precludes searching for interactions and creates difficulties for replicating studies.\n\nResultsIn this work, we develop a new statistical approach to address these issues that leverages genetic ancestry in admixed populations. We applied our method to gene expression and methylation data from African American and Latino admixed individuals respectively, identifying nine interactions that were significant at p < 5x10-8, we show that two of the interactions in methylation data replicate, and the remaining six are significantly enriched for low p-values (p < 1.8x10-6).\n\nConclusionWe show that genetic ancestry can be a useful proxy for unknown and unmeasured covariates in the search for interaction effects. These results have important implications for our understanding of the genetic architecture of complex traits.

Genetics

Fast and accurate construction of confidence intervals for heritability

Estimation of heritability is fundamental in genetic studies. In recent years, heritability estimation using linear mixed models (LMMs) has gained popularity, because these estimates can be obtained from unrelated individuals collected in genome wide association studies. Typically, heritability estimation under LMMs uses either the maximum likelihood (ML) or the restricted maximum likelihood (REML) approach. Existing methods for the construction of confidence intervals and estimators of standard errors for both ML and REML rely on asymptotic properties. However, these assumptions are often violated due to the bounded parameter space, statistical dependencies, and limited sample size, leading to biased estimates, and inflated or deflated confidence intervals. Here, we show that often the probability that the genetic component is estimated as zero is high even when the true heritability is bounded away from zero, emphasizing the need for accurate confidence intervals. We further show that the estimation of confidence intervals by state-of-the-art methods is highly inaccurate, especially when the true heritability is either relatively low or relatively high. Such biases are present, for example, in estimates of heritability of gene expression in the GTEx study, and of lipid profiles in the LURIC study. We propose a computationally efficient method, Accurate LMM-Based confidence Intervals (ALBI), for the estimation of the distribution of the heritability estimator, and for the construction of accurate confidence intervals. Our method can be used as an add-on to existing methods for heritability and variance components estimation, such as GCTA, FaST-LMM, GEMMA, or EMMA. ALBI is available at http://www.cs.tau.ac.il/~heran/cozygene/software/albi.html.

Genetics

Recycler: an algorithm for detecting plasmids from de novo assembly graphs

Plasmids are central contributors to microbial evolution and genome innovation. Recently, they have been found to have important roles in antibiotic resistance and in affecting production of metabolites used in industrial and agricultural applications. However, their characterization through deep sequencing remains challenging, in spite of rapid drops in cost and throughput increases for sequencing. Here, we attempt to ameliorate this situation by introducing a new plasmid-specific assembly algorithm, leveraging assembly graphs provided by a conventional de novo assembler and alignments of paired- end reads to assembled graph nodes. We introduce the first tool for this task, called Recycler, and demonstrate its merits in comparison with extant approaches. We show that Recycler greatly increases the number of true plasmids recovered while remaining highly accurate. On simulated plasmidomes, Recycler recovered 5-14% more true plasmids compared to the best extant method with overall precision of about 90%. We validated these results in silico on real data, as well as in vitro by PCR validation performed on a subset of Recyclers predictions on different data types. All 12 of Recyclers outputs on isolate samples matched known plasmids or phages, and had alignments having at least 97% identity over at least 99% of the reported reference sequence lengths. For the two E. Coli strains examined, most known plasmid sequences were recovered, while in both cases additional plasmids only known to be present in different hosts were found. Recycler also generated plasmids in high agreement with known annotation on real plasmidome data. Moreover, in PCR validations performed on 77 sequences, Recycler showed mean accuracy of 89% across all data types - isolate, microbiome, and plasmidome. Recycler is available at http://github.com/Shamir-Lab/Recycler

Genomics

The genetics of Bene Israel from India reveals both substantial Jewish and Indian ancestry

The Bene Israel Jewish community from West India is a unique population whose history before the 18th century remains largely unknown. Bene Israel members consider themselves as descendants of Jews, yet the identity of Jewish ancestors and their arrival time to India are unknown, with speculations on arrival time varying between the 8th century BCE and the 6th century CE. Here, we characterize the genetic history of Bene Israel by collecting and genotyping 18 Bene Israel individuals. Combining with 486 individuals from 41 other Jewish, Indian and Pakistani populations, and additional individuals from worldwide populations, we conducted comprehensive genome-wide analyses based on FST, principal component analysis, ADMIXTURE, identity-by-descent sharing, admixture linkage disequilibrium decay, haplotype sharing and allele sharing autocorrelation decay, as well as contrasted patterns between the X chromosome and the autosomes. The genetics of Bene Israel individuals resemble local Indian populations, while at the same time constituting a clearly separated and unique population in India. They are unique among Indian and Pakistani populations we analyzed in sharing considerable genetic ancestry with other Jewish populations. Putting together the results from all analyses point to Bene Israel being an admixed population with both Jewish and Indian ancestry, with the genetic contribution of each of these ancestral populations being substantial. The admixture took place in the last millennium, about 19-33 generations ago. It involved Middle-Eastern Jews and was sex-biased, with more male Jewish and local female contribution. It was followed by a population bottleneck and high endogamy, which can lead to increased prevalence of recessive diseases in this population. This study provides an example of how genetic analysis advances our knowledge of human history in cases where other disciplines lack the relevant data to do so.

Genetics