bioRxiv ScienceSearch

Biology subjects

Voight, B. F.

Publications and source records attributed to Voight, B. F..

5 recordsLinked to original sources

Signals of variation in human mutation rate at multiple levels of sequence context

Our understanding of mutation rate helps us build evolutionary models and make sense of genetic variation. Recent work indicates that the frequencies of specific mutation types have been elevated in Europe, and that many more, subtler signatures of global polymorphism variation may yet remain unidentified. Here, we present an analysis of the 1,000 Genomes Project (phase 3), suggesting additional putative signatures of mutation rate variation across populations and the extent to which they are shaped by local sequence context. First, we compiled a list of the most significantly variable polymorphism types in a cross-continental statistical test. Clustering polymorphisms together, we observed four sets of substitution types that showed similar trends of relative mutation rate across populations, and describe the patterns of these mutational clusters among continental groups. For the majority of these signatures, we found that a single flanking base pair of sequence context was sufficient to determine the majority of enrichment or depletion of a mutation type. However, local genetic context up to 2-3 base pairs away contributes additional variability, and helps to interpret a previously noted enrichment of certain polymorphism types in some East Asian groups. Building our understanding of mutation rate in this way can help us to construct more accurate evolutionary models and better understand the mechanisms that underlie genetic change.

genomics

Bivariate GWAS scan identifies six novel loci associated with lipid levels and coronary artery disease

BackgroundPlasma lipid levels are heritable and genetically associated with risk of coronary artery disease (CAD). However, genome-wide association studies (GWAS) routinely analyze these traits independently of one another. Joint GWAS for two related phenotypes can lead to a higher-powered analysis to detect variants contributing to both traits.\n\nMethods and ResultsWe performed a bivariate GWAS to discover novel loci associated with heart disease, using a CAD Meta-Analysis (122,733 cases and 424,528 controls), and lipid traits, using data from the Global Lipid Genetics Consortium (188,577 subjects). We identified six previously unreported loci at genome-wide significance (P < 5 x 10-8), three which were associated with Triglycerides and CAD, two which were associated with LDL cholesterol and CAD, and one associated with Total Cholesterol and CAD. At several of our loci, the GWAS signals jointly localize with genetic variants associated with expression level changes for one or more neighboring genes, indicating that these loci may be affecting disease risk through regulatory activity.\n\nConclusionsWe discovered six novel variants individually associated with both lipids and coronary artery disease.

genomics

Multiplexed targeted resequencing identifies coding and regulatory variation underlying phenotypic extremes of HDL-cholesterol in humans

Genome-wide association studies have uncovered common variants at many loci influencing human complex traits and diseases, such as high-density lipoprotein cholesterol (HDL-C). However, the contribution of the identified genes is difficult to ascertain from current efforts interrogating common variants with small effects. Thus, there is a pressing need for scalable, cost-effective strategies for uncovering causal variants, many of which may be rare and noncoding. Here, we used a multiplexed inversion probe (MIP) target capture approach to resequence both coding and regulatory regions at seven HDL-C associated loci in 797 individuals with extremely high HDL-C vs. 735 low-to-normal HDL-C controls. Our targets included protein-coding regions of GALNT2, APOA5, APOC3, SCARB1, CCDC92, ZNF664, CETP, and LIPG (>9 kb), and proximate noncoding regulatory features (>42 kb). Exome-wide genotyping in 1,114 of the 1,532 participants yielded a >90% genotyping concordance rate with MIP-identified variants in ~90% of participants. This approach rediscovered nearly all established GWAS associations in GALNT2, CETP, and LIPG loci with significant and concordant associations with HDL-C from our phenotypic-extremes design at 0.1% of the sample size of lipid GWAS studies. In addition, we identified a novel, rare, CETP noncoding variant enriched in the extreme high HDL-C group (P<0.01, Score Test). Our targeted resequencing of individuals at the HDL-C phenotypic extremes offers a novel, efficient, and cost-effective approach for identifying rare coding and noncoding variation differences in extreme phenotypes and supports the rationale for applying this methodology to uncover rare variation--particularly non-coding variation--underlying myriad complex traits.

genomics

Detecting Long-term Balancing Selection using Allele Frequency Correlation

Balancing selection occurs when multiple alleles are maintained in a population, which can result in their preservation over long evolutionary time periods. A characteristic signature of this long-term balancing selection is an excess number of intermediate frequency polymorphisms near the balanced variant. However, the expected distribution of allele frequencies at these loci has not been extensively detailed, and therefore existing summary statistic methods do not explicitly take it into account. Using simulations, we show that new mutations which arise in close proximity to a site targeted by balancing selection accumulate at frequencies nearly identical to that of the balanced allele. In order to scan the genome for balancing selection, we propose a new summary statistic, {beta}, which detects these clusters of alleles at similar frequencies. Simulation studies show that compared to existing summary statistics, our measure has improved power to detect balancing selection, and is reasonably powered in non-equilibrium demographic models or when recombination or mutation rate varies. We compute {beta} on 1000 Genomes Project data to identify lo ci potentially subjected to long-term balancing selection in humans. We report two balanced haplotypes - localized to the genes WFS1 and CADM2 - that are strongly linked to association signals for complex traits. Our approach is computationally efficient and applicable to species that lack appropriate outgroup sequences, allowing for well-powered analysis of selection in the wide variety of species for which population data are rapidly being generated.

evolutionary biology

Patterns of shared signatures of recent positive selection across human populations

Scans for positive selection in human populations have identified hundreds of sites across the genome with evidence of recent adaptation. These signatures often overlap across populations, but the question of how often these overlaps represent a single ancestral event remains unresolved. If a single positive selection event spread across many populations, the same sweeping haplotype should appear in each population and the selective pressure could be common across diverse populations and environments. Identifying such shared selective events would be of fundamental interest, pointing to genomic loci and human traits important in recent history across the globe. Additionally, genomic annotations that recently became available could help attach these signatures to a potential gene and molecular phenotype that may have been selected across multiple populations. We performed a scan for positive selection using the integrated haplotype score on 20 populations, and compared sweeping haplotypes using the haplotype-clustering capability of fastPHASE to create a catalog of shared and unshared overlapping selective sweeps in these populations. Using additional genomic annotations, we connect these multi-population sweep overlaps with potential biological mechanisms at several loci, including potential new sites of adaptive introgression, the glycophorin locus associated with malarial resistance, and the alcohol dehydrogenase cluster associated with alcohol dependency.

genetics