bioRxiv ScienceSearch

Biology subjects

Rasmus Nielsen

Publications and source records attributed to Rasmus Nielsen.

11 recordsLinked to original sources

Estimating the timing of multiple admixture events using 3-locus Linkage Disequilibrium

Estimating admixture histories is crucial for understanding the genetic diversity we see in present-day populations. Allele frequency or phylogeny-based methods are excellent for inferring the existence of admixture or its proportions. However, to estimate admixture times, spatial information from admixed chromosomes of local ancestry or the decay of admixture linkage disequilibrium (ALD) is used. One popular method, implemented in the programs ALDER and ROLLOFF, uses two-locus ALD to infer the time of a single admixture event, but is only able to estimate the time of the most recent admixture event based on this summary statistic. To address this limitation, we derive analytical expressions for the expected ALD in a three-locus system and provide a new statistical method based on these results that is able to resolve more complicated admixture histories. Using simulations, we evaluate the performance of this method on a range of different admixture histories. As an example, we apply the method to the Colombian and Mexican samples from the 1000 Genomes project. The implementation of our method is available at https://github.com/Genomics-HSE/LaNeta. Author summaryWe establish a theoretical framework to model 3-locus admixture linkage disequilibrium of an admixed population taking into account the effects of genetic drift, migration and recombination. The theory is used to develop a method for estimating the times of multiple admixtures events. We demonstrate the accuracy of the method on simulated data and we apply it to previously published data from Mexican and Columbian populations to explore the complex history of American populations in the post-Columbian period.

Bioinformatics

Ohana, a tool set for population genetic analyses of admixture components

MotivationStructure methods are highly used population genetic methods for classifying individuals in a sample fractionally into discrete ancestry components.\n\nContributionWe introduce a new optimization algorithm of the classical Structure model in a maximum likelihood framework. Using analyses of real data we show that the new optimization algorithm finds higher likelihood values than the state-of-the-art method in the same computational time. We also present a new method for estimating population trees from ancestry components using a Gaussian approximation. Using coalescence simulations modeling populations evolving in a tree-like fashion, we explore the adequacy of the Structure model and the Gaussian assumption for identifying ancestry components correctly and for inferring the correct tree. In most cases, ancestry components are inferred correctly, although sample sizes and times since admixture can influence the inferences. Similarly, the popular Gaussian approximation tends to perform poorly when branch lengths are long, although the tree topology is correctly inferred in all scenarios explored. The new methods are implemented together with appropriate visualization tools in the computer package Ohana.\n\nAvailabilityOhana is publicly available at https://github.com/jade-cheng/ohana. Besides its source code and installation instructions, we also provide example workflows in the project wiki site.\n\nContactjade.cheng@birc.au.dk

Bioinformatics

A hidden Markov model approach for simultaneously estimating local ancestry and admixture time using next generation sequence data in samples of arbitrary ploidy

Admixture--the mixing of genomes from divergent populations--is increasingly appreciated as a central process in evolution. To characterize and quantify patterns of admixture across the genome, a number of methods have been developed for local ancestry inference. However, existing approaches have a number of shortcomings. First, all local ancestry inference methods require some prior assumption about the expected ancestry tract lengths. Second, existing methods generally require genotypes, which is not feasible to obtain for many next-generation sequencing projects. Third, many methods assume samples are diploid, however a wide variety of sequencing applications will fail to meet this assumption. To address these issues, we introduce a novel hidden Markov model for estimating local ancestry that models the read pileup data, rather than genotypes, is generalized to arbitrary ploidy, and can estimate the time since admixture during local ancestry inference. We demonstrate that our method can simultaneously estimate the time since admixture and local ancestry with good accuracy, and that it performs well on samples of high ploidy--i.e. 100 or more chromosomes. As this method is very general, we expect it will be useful for local ancestry inference in a wider variety of populations than what previously has been possible. We then applied our method to pooled sequencing data derived from populations of Drosophila melanogaster on an ancestry cline on the east coast of North America. We find that regions of local recombination rates are negatively correlated with the proportion of African ancestry, suggesting that selection against foreign ancestry is the least efficient in low recombination regions. Finally we show that clinal outlier loci are enriched for genes associated with gene regulatory functions, consistent with a role of regulatory evolution in ecological adaptation of admixed D. melanogaster populations. Our results illustrate the potential of local ancestry inference for elucidating fundamental evolutionary processes.\n\nAuthor SummaryWhen divergent populations hybridize, their offspring obtain portions of their genomes from each parent population. Although the average ancestry proportion in each descendant is equal to the proportion of ancestors from each of the ancestral populations, the contribution of each ancestry type is variable across the genome. Estimating local ancestry within admixed individuals is a fundamental goal for evolutionary genetics, and here we develop a method for doing this that circumvents many of the problems associated with existing methods. Briefly, our method can use short read data, rather than genotypes and can be applied to samples with any number of chromosomes. Furthermore, our method simultaneously estimates local ancestry and the number of generations since admixture--the time that the two ancestral populations first encountered each other. Finally, in applying our method to data from an admixture zone between ancestral populations of Drosophila melanogaster, we find many lines of evidence consistent with natural selection operating to against the introduction of foreign ancestry into populations of one predominant ancestry type. Because of the generality of this method, we expect that it will be useful for a wide variety of existing and ongoing research projects.

Evolutionary Biology

Archaic adaptive introgression in TBX15/WARS2

A recent study conducted the first genome-wide scan for selection in Inuit from Greenland using SNP chip data. Here, we report that selection in the region with the second most extreme signal of positive selection in Greenlandic Inuit favored a deeply divergent haplotype that is closely related to the sequence in the Denisovan genome, and was likely introgressed from an archaic population. The region contains two genes, WARS2 and TBX15, and has previously been associated with adipose tissue differentiation and body-fat distribution in humans. We show that the adaptively introgressed allele has been under selection in a much larger geographic region than just Greenland. Furthermore, it is associated with changes in expression of WARS2 and TBX15 in multiple tissues including the adrenal gland and subcutaneous adipose tissue, and with regional DNA methylation changes in TBX15.

Evolutionary Biology

The Genetic Cost of Neanderthal Introgression

Approximately 2-4% of genetic material in human populations outside Africa is derived from Neanderthals who interbred with anatomically modern humans. Recent studies have shown that this Neanderthal DNA is depleted around functional genomic regions; this has been suggested to be a consequence of harmful epistatic interactions between human and Neanderthal alleles. However, using published estimates of Neanderthal inbreeding and the distribution of mutational fitness effects, we infer that Neanderthals had at least 40% lower fitness than humans on average; this increased load predicts the reduction in Neanderthal introgression around genes without the need to invoke epistasis. We also predict a residual Neanderthal mutational load in non-Africans, leading to a fitness reduction of at least 0.5%. This effect of Neanderthal admixture has been left out of previous debate on mutation load differences between Africans and non-Africans. We also show that if many deleterious mutations are recessive, the Neanderthal admixture fraction could increase over time due to the protective effect of Neanderthal haplotypes against deleterious alleles that arose recently in the human population. This might partially explain why so many organisms retain gene flow from other species and appear to derive adaptive benefits from introgression.

Evolutionary Biology

Detecting recent selective sweeps while controlling for mutation rate and background selection

A composite likelihood ratio test implemented in the program SweepFinder is a commonly used method for scanning a genome for recent selective sweeps. SweepFinder uses information on the spatial pattern of the site frequency spectrum (SFS) around the selected locus. To avoid confounding effects of background selection and variation in the mutation process along the genome, the method is typically applied only to sites that are variable within species. However, the power to detect and localize selective sweeps can be greatly improved if invariable sites are also included in the analysis. In the spirit of a Hudson-Kreitman-Aguade test, we suggest to add fixed differences relative to an outgroup to account for variation in mutation rate, thereby facilitating more robust and powerful analyses. We also develop a method for including background selection modeled as a local reduction in the effective population size. Using simulations we show that these advances lead to a gain in power while maintaining robustness to mutation rate variation. Furthermore, the new method also provides more precise localization of the causative mutation than methods using the spatial pattern of segregating sites alone.

Evolutionary Biology

The origin and evolution of maize in the American Southwest

Maize offers an ideal system through which to demonstrate the potential of ancient population genomic techniques for reconstructing the evolution and spread of domesticates. The diffusion of maize from Mexico into the North American Southwest (SW) remains contentious with the available evidence being restricted to morphological studies of ancient maize plant material. We captured 1 Mb of nuclear DNA from 32 archaeological maize samples spanning 6000 years and compared them with modern landraces including those from the Mexican West coast and highlands. We found that the initial diffusion of domesticated maize into the SW is likely to have occurred through a highland route. However, by 2000 years ago a Pacific coastal corridor was also being used. Furthermore, we could distinguish between genes that were selected for early during domestication (such as zagl1 involved in shattering) from genes that changed in the SW context (e.g. related to sugar content and adaptation to drought) likely as a response to the local arid environment and new cultural uses of maize.

Evolutionary Biology

Fitting the Balding-Nichols model to forensic databases

Large forensic databases provide an opportunity to compare observed empirical rates of genotype matching with those expected under forensic genetic models. A number of researchers have taken advantage of this opportunity to validate some forensic genetic approaches, particularly to ensure that estimated rates of genotype matching between unrelated individuals are indeed slight overestimates of those observed. However, these studies have also revealed systematic error trends in genotype probability estimates. In this analysis, we investigate these error trends and show how the specific implementation of the Balding-Nichols model must be considered when applied to database-wide matching. Specifically, we show that in addition to accounting for increased allelic matching between individuals with recent shared ancestry, studies must account for relatively decreased allelic matching between individuals with more ancient shared ancestry.

Genetics

Reticulate speciation and adaptive introgression in the Anopheles gambiae species complex

Anopheles gambiae, the primary vector of human malaria in sub-Saharan Africa, exists as a series of ecologically specialized subgroups that are phylogenetically nested within the Anopheles gambiae species complex. These species and subgroups exhibit varying degrees of reproductive isolation, sometimes recognized as distinct subspecies. We have sequenced 32 complete genomes from field-captured individuals of Anopheles gambiae, Anopheles gambiae M form (recently named A. coluzzii), sister species A. arabiensis, and the recently discovered \"GOUNDRY\" subgroup of A. gambiae that is highly susceptible to Plasmodium. Amidst a backdrop of strong reproductive isolation and adaptive differentiation, we find evidence for introgression of autosomal chromosomal regions among species and subgroups, some of which have facilitated adaptation. The X chromosome, however, is strongly differentiated among all species and subgroups, pointing to a disproportionately large effect of X chromosome genes in driving speciation among anophelines. Strikingly, we find that autosomal introgression has occurred from contemporary hybridization among A. gambiae and A. arabiensis despite strong divergence ({bsim}5x higher than autosomal divergence) and isolation on the X chromosome. We find a large region of the X chromosome that has swept to fixation in the GOUNDRY subgroup within the last 100 years, which may be an inversion that serves as a partial barrier to contemporary gene flow. We show that speciation with gene flow results in genomic mosaicism of divergence and introgression. Such a reticulate gene pool connecting vector species and subgroups across the speciation continuum has important implications for malaria control efforts.\n\nAuthor SummarySubdivision of species into ecological specialized subgroups allows organisms to access a wider variety of environments and sometimes leads to the formation of species complexes. Adaptation to distinct environments tends to result in differentiation among closely related subgroups, although hybridization can facilitate sharing of globally adaptive alleles. Here, we show that differentiation and hybridization have acted in parallel in a species complex of Anopheles mosquitoes that vector human malaria. In particular, we show that extensive adaptive differentiation and partial reproductive isolation has led to genomic differentiation among mosquito species and subgroups, especially on the X chromosome. However, we also find evidence for exchange of genes on the autosomes that has provided the raw material for recent rapid adaptation. For example, we show that A. arabiensis has shared a mutation conferring insecticide resistance with two subgroups of A. gambiae within the last 60 years, illustrating the fluid nature of species boundaries among even more advanced species pairs. Our results underscore the expected challenges in deploying vector-based disease control strategies since many of the worlds most devastating human pathogens are transmitted by arthropod species complexes.

Evolutionary Biology

Understanding Admixture Fractions

Estimation of admixture fractions has become one of the most commonly used computational tools in population genomics. However, there is remarkably little population genetic theory on their statistical properties. We develop theoretical results that can accurately predict means and variances of admixture proportions within a population using models with recombination and genetic drift. Based on established theory on measures of multilocus disequilibrium, we show that there is a set of recurrence relations that can be used to derive expectations for higher moments of the admixture fraction distribution. We obtain closed form solutions for some special cases. Using these results, we develop a method for estimating admixture parameters from estimated admixture proportion obtained from programs such as Structure or Admixture. We apply this method to HapMap data and find that the population history of African Americans, as expected, is not best explained by a single admixture event between people of European and African ancestry. A model of constant gene flow for the past 11 generations until 2 generations ago gives a better fit.

Evolutionary Biology

Phylogenetic ANOVA: The Expression Variance and Evolution (EVE) model for quantitative trait evolution

A number of methods have been developed for modeling the evolution of a quantitative trait on a phylogeny. These methods have received renewed interest in the context of genome-wide studies of gene expression, in which the expression levels of many genes can be modeled as quantitative traits. We here develop a new method for joint analyses of quantitative traits within and between-species, the Expression Variance and Evolution (EVE) model. The model parameterizes the ratio of population to evolutionary expression variance, facilitating a wide variety of analyses, including a test for lineage-specific shifts in expression level, and a phylogenetic ANOVA that can detect genes with increased or decreased ratios of expression divergence to diversity, analogous to the famous HKA test used to detect selection at the DNA level. We use simulations to explore the properties of these tests under a variety of circumstances and show that the phylogenetic ANOVA is more accurate than the standard ANOVA (no accounting for phylogeny) sometimes used in transcriptomics. We then apply the EVE model to a mammalian phylogeny of 15 species typed for expression levels in liver tissue. We identify genes with high expression divergence between-species as candidates for expression level adaptation, and genes with high expression diversity within-species as candidates for expression level conservation and/or plasticity. Using the test for lineage-specific expression shifts, we identify several candidate genes for expression level adaptation on the catarrhine and human lineages, including genes putatively related to dietary changes in humans. We compare these results to those reported previously using a model which ignores expression variance within-species, uncovering important differences in performance. We demonstrate the necessity for a phylogenetic model in comparative expression studies and show the utility of the EVE model to detect expression divergence, diversity, and branch-specific shifts.

Evolutionary Biology