bioRxiv ScienceSearch

bioRxiv · 10.1101/024042

Inference and analysis of population structure using genetic data and network theory

Abstract

Clustering individuals to subpopulations based on genetic data has become commonplace in many genetic studies. Inference of population structure is most often done by applying model-based approaches, aided by visualization using distance-based approaches such as multidimensional scaling. While existing distance-based approaches suffer from lack of statistical rigor, model-based approaches entail assumptions of prior conditions such as that the subpopulations are at Hardy-Weinberg equilibria. Here we present a distance-based approach for inference of population structure using genetic data by defining population structure using network theory terminology and methods. A network is constructed from a pairwise genetic-similarity matrix of all sampled individuals. The community partition, a partition of a network to dense subgraphs, is equated with population structure, a partition of the population to genetically related groups. Community detection algorithms are used to partition the network into communities, interpreted as a partition of the population to subpopulations. The statistical significance of the structure can be estimated by using permutation tests to evaluate the significance of the partitions modularity, a network theory measure indicating the quality of community partitions. In order to further characterize population structure, a new measure of the Strength of Association (SA) for an individual to its assigned community is presented. The Strength of Association Distribution (SAD) of the communities is analyzed to provide additional population structure characteristics, such as the relative amount of gene flow experienced by the different subpopulations and identification of hybrid individuals. Human genetic data and simulations are used to demonstrate the applicability of the analyses. The approach presented here provides a novel, computationally efficient, model-free method for inference of population structure which does not entail assumption of prior conditions. The method is implemented in the software NetStruct, available at https://github.com/GiliG/NetStruct.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Gili Greenbaum, Alan R. Templeton, Shirli Bar-David. 2015-08-06. Inference and analysis of population structure using genetic data and network theory. https://doi.org/10.1101/024042

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Rapid evolution of primate type 2 immune response factors linked to asthma susceptibility

Host immunity pathways evolve rapidly in response to antagonism by pathogens. Microbial infections can also trigger excessive inflammation that contributes to diverse autoimmune disorders including asthma, lupus, diabetes, and arthritis. Definitive links between immune system evolution and human autoimmune disease remain unclear. Here we provide evidence that several components of the type 2 immune response pathway have been subject to recurrent positive selection in the primate lineage. Notably, rapid evolution of the central immune regulator IL13 corresponds to a polymorphism linked to asthma susceptibility in humans. We also find evidence of accelerated amino acid substitutions as well as repeated gene gain and loss events among eosinophil granule proteins, which act as toxic antimicrobial effectors that promote asthma pathology by damaging airway tissues. These results support the hypothesis that evolutionary conflicts with pathogens promote tradeoffs for increasingly robust immune responses during animal evolution. Our findings are also consistent with the view that natural selection has contributed to the spread of autoimmune disease alleles in humans.

Genetics

Single Cell Expression Data Reveal Human Genes that Escape X-Chromosome Inactivation

Sex chromosomes pose an inherent genetic imbalance between genders. In mammals, one of the females X-chromosomes undergoes inactivation (Xi). Indirect measurements estimate that about 20% of Xi genes completely or partially escape inactivation. The identity of these escapee genes and their propensity to escape inactivation remain unsolved. A direct method for identifying escapees was applied by quantifying differential allelic expression from single cells. RNA-Seq fragments were assigned to informative SNPs which were labeled by the appropriate parental haplotype. This method was applied for measuring allelic specific expression from Chromosome-X (ChrX) and an autosomal chromosome as a control. We applied the protocol for measuring biallelic expression from ChrX to 104 primary fibroblasts. Out of 215 genes that were considered, only 13 genes (6%) were associated with biallelic expression. The sensitivity of escapees' identification was increased by combining SNP mapping for parental diploid genomes together with RNA-Seq from clonal single cells (25 lymphoblasts). Using complementary protocols, referred to as strict and relaxed, we confidently identified 25 and 31escapee genes, respectively. When pooled versions of 30 and 100 cells were used, <50% of these genes were revealed. We assessed the generality of our protocols in view of an escapee catalog compiled from indirect methods. The overlap between the escapee catalog and the genes list from this study is statistically significant (P-value of E-07). We conclude that single cells expression data are instrumental for studying X-inactivation with an improved sensitivity. Finally, our results support the emerging notion of the non-deterministic nature of genes that escape X-chromosome inactivation.

Genetics

Frequency of mosaicism points towards mutation-prone early cleavage cell divisions.

It has recently become possible to directly estimate the germ-line de novo mutation (dnm) rate by sequencing the whole genome of father-mother-offspring trios, and this has been conducted in human1-5, chimpanzee6, mice7, birds8 and fish9. In these studies dnms are typically defined as variants that are heterozygous in the offspring while being absent in both parents. They are assumed to have occurred in the germ-line of one of the parents and to have been transmitted to the offspring via the sperm cell or oocyte. This definition assumes that detectable mosaicism in the parent in which the mutation occurred is negligible. However, instances of detectable mosaicism or premeiotic clusters are well documented in humans and other organisms, including ruminants10-12. We herein take advantage of cattle pedigrees to show that as much as [~]30% to [~]50% of dnms present in a gamete may occur during the early cleavage cell divisions in males and females, respectively, resulting in frequent detectable mosaicism and a high rate of sharing of multiple dnms between siblings. This should be taken into account to accurately estimate the mutation rate in cattle and other species.

Genetics