bioRxiv ScienceSearch

Biology subjects

Omri Tal

Publications and source records attributed to Omri Tal.

2 recordsLinked to original sources

From Typical Sequences to Typical Genotypes

We demonstrate an application of a core notion of information theory, that of typical sequences and their related properties, to analysis of population genetic data. Based on the asymptotic equipartition property (AEP) for non-stationary discrete-time sources producing independent symbols, we introduce the concepts of typical genotypes and population entropy rate and cross entropy rate. We analyze three perspectives on typical genotypes: a set perspective on the interplay of typical sets of genotypes from two populations, a geometric perspective on their structure in high dimensional space, and a statistical learning perspective on the prospects of constructing typical-set based classifiers. In particular, we show that such classifiers have a surprising resilience to noise originating from small population samples, and highlight the potential for further links between inference and communication.

Genetics

On the Apportionment of Population Structure

Measures of population differentiation, such as FST, are traditionally derived from the partition of diversity within and between populations. However, the emergence of population clusters from multilocus analysis is a function of genetic structure (departures from panmixia) rather than of diversity. If the populations are close to panmixia, slight differences between the mean pairwise distance within and between populations (low FST) can manifest as strong separation between the populations, thus population clusters are often evident even when the vast majority of diversity is partitioned within populations rather than between them. For any given FST value, clusters can be tighter (more panmictic) or looser (more stratified), and in this respect higher FST does not always imply stronger differentiation. In this study we propose a measure for the partition of structure, denoted EST, which is more consistent with results from clustering schemes. Crucially, our measure is based on a statistic of the data that is a good measure of internal structure, mimicking the information extracted by unsupervised clustering or dimensionality reduction schemes. To assess the utility of our metric, we ranked various human (HGDP) population pairs based on FST and EST and found substantial differences in ranking order. In some cases examined, most notably among isolated Amazonian tribes, EST ranking seems more consistent with demographic, phylogeographic and linguistic measures of classification compared to FST. Thus, EST may at times outperform FST in identifying evolutionary significant differentiation.

Evolutionary Biology