bioRxiv ScienceSearch

Biology subjects

Saharon Rosset

Publications and source records attributed to Saharon Rosset.

2 recordsLinked to original sources

On the Apportionment of Population Structure

Measures of population differentiation, such as FST, are traditionally derived from the partition of diversity within and between populations. However, the emergence of population clusters from multilocus analysis is a function of genetic structure (departures from panmixia) rather than of diversity. If the populations are close to panmixia, slight differences between the mean pairwise distance within and between populations (low FST) can manifest as strong separation between the populations, thus population clusters are often evident even when the vast majority of diversity is partitioned within populations rather than between them. For any given FST value, clusters can be tighter (more panmictic) or looser (more stratified), and in this respect higher FST does not always imply stronger differentiation. In this study we propose a measure for the partition of structure, denoted EST, which is more consistent with results from clustering schemes. Crucially, our measure is based on a statistic of the data that is a good measure of internal structure, mimicking the information extracted by unsupervised clustering or dimensionality reduction schemes. To assess the utility of our metric, we ranked various human (HGDP) population pairs based on FST and EST and found substantial differences in ranking order. In some cases examined, most notably among isolated Amazonian tribes, EST ranking seems more consistent with demographic, phylogeographic and linguistic measures of classification compared to FST. Thus, EST may at times outperform FST in identifying evolutionary significant differentiation.

Evolutionary Biology

Fast and accurate construction of confidence intervals for heritability

Estimation of heritability is fundamental in genetic studies. In recent years, heritability estimation using linear mixed models (LMMs) has gained popularity, because these estimates can be obtained from unrelated individuals collected in genome wide association studies. Typically, heritability estimation under LMMs uses either the maximum likelihood (ML) or the restricted maximum likelihood (REML) approach. Existing methods for the construction of confidence intervals and estimators of standard errors for both ML and REML rely on asymptotic properties. However, these assumptions are often violated due to the bounded parameter space, statistical dependencies, and limited sample size, leading to biased estimates, and inflated or deflated confidence intervals. Here, we show that often the probability that the genetic component is estimated as zero is high even when the true heritability is bounded away from zero, emphasizing the need for accurate confidence intervals. We further show that the estimation of confidence intervals by state-of-the-art methods is highly inaccurate, especially when the true heritability is either relatively low or relatively high. Such biases are present, for example, in estimates of heritability of gene expression in the GTEx study, and of lipid profiles in the LURIC study. We propose a computationally efficient method, Accurate LMM-Based confidence Intervals (ALBI), for the estimation of the distribution of the heritability estimator, and for the construction of accurate confidence intervals. Our method can be used as an add-on to existing methods for heritability and variance components estimation, such as GCTA, FaST-LMM, GEMMA, or EMMA. ALBI is available at http://www.cs.tau.ac.il/~heran/cozygene/software/albi.html.

Genetics