bioRxiv ScienceSearch

Biology subjects

Petkova, D.

Publications and source records attributed to Petkova, D..

3 recordsLinked to original sources

Estimating recent migration and population size surfaces

In many species a fundamental feature of genetic diversity is that genetic similarity decays with geographic distance; however, this relationship is often complex, and may vary across space and time. Methods to uncover and visualize such relationships have widespread use for analyses in molecular ecology, conservation genetics, evolutionary genetics, and human genetics. While several frameworks exist, a promising approach is to infer maps of how migration rates vary across geographic space. Such maps could, in principle, be estimated across time to reveal the full complexity of population histories. Here, we take a step in this direction: we present a method to infer separate maps of population sizes and migration rates for different time periods from a matrix of genetic similarity between every pair of individuals. Specifically, genetic similarity is measured by counting the number of long segments of haplotype sharing (also known as identity-by-descent tracts). By varying the length of these segments we obtain parameter estimates for qualitatively different time periods. Using simulations, we show that the method can reveal time-varying migration rates and population sizes, including changes that are not detectable when ignoring haplotypic structure. We apply the method to a dataset of contemporary European individuals (POPRES), and provide an integrated analysis of recent population structure and growth over the last ~3,000 years in Europe. Software implementing the methods is available at https://github.com/halasadi/MAPS.

bioinformatics

Genetic landscapes reveal how human genetic diversity aligns with geography

Geographic patterns in human genetic diversity carry footprints of population history1,2 and provide insights for genetic medicine and its application across human populations3,4. Summarizing and visually representing these patterns of diversity has been a persistent goal for human geneticists5-10, and has revealed that genetic differentiation is frequently correlated with geographic distance. However, most analytical methods to represent population structure11-15 do not incorporate geography directly, and it must be considered post hoc alongside a visual summary. Here, we use a recently developed spatially explicit method to estimate \"effective migration\" surfaces to visualize how human genetic diversity is geographically structured (the EEMS method16). The resulting surfaces are \"rugged\", which indicates the relationship between genetic and geographic distance is heterogenous and distorted as a rule. Most prominently, topographic and marine features regularly align with increased genetic differentiation (e.g. the Sahara desert, Mediterranean Sea or Himalaya at large scales; the Adriatic, interisland straits in near Oceania at smaller scales). In other cases, the locations of historical migrations and boundaries of language families align with migration features. These results provide visualizations of human genetic diversity that reveal local patterns of differentiation in detail and emphasize that while genetic similarity generally decays with geographic distance, there have regularly been factors that subtly distort the underlying relationship across space observed today. The fine-scale population structure depicted here is relevant to understanding complex processes of human population history and may provide insights for geographic patterning in rare variants and heritable disease risk.

evolutionary biology

Genome-wide genetic data on ~500,000 UK Biobank participants

The UK Biobank project is a large prospective cohort study of ~500,000 individuals from across the United Kingdom, aged between 40-69 at recruitment. A rich variety of phenotypic and health-related information is available on each participant, making the resource unprecedented in its size and scope. Here we describe the genome-wide genotype data (~805,000 markers) collected on all individuals in the cohort and its quality control procedures. Genotype data on this scale offers novel opportunities for assessing quality issues, although the wide range of ancestries of the individuals in the cohort also creates particular challenges. We also conducted a set of analyses that reveal properties of the genetic data - such as population structure and relatedness - that can be important for downstream analyses. In addition, we phased and imputed genotypes into the dataset, using computationally efficient methods combined with the Haplotype Reference Consortium (HRC) and UK10K haplotype resource. This increases the number of testable variants by over 100-fold to ~96 million variants. We also imputed classical allelic variation at 11 human leukocyte antigen (HLA) genes, and as a quality control check of this imputation, we replicate signals of known associations between HLA alleles and many common diseases. We describe tools that allow efficient genome-wide association studies (GWAS) of multiple traits and fast phenome-wide association studies (PheWAS), which work together with a new compressed file format that has been used to distribute the dataset. As a further check of the genotyped and imputed datasets, we performed a test-case genome-wide association scan on a well-studied human trait, standing height.

genetics