bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Genetic Allee effects and their interaction with ecological Allee effects

SummaryO_LIIt is now widely accepted that genetic processes such as inbreeding depression and loss of genetic variation can increase the extinction risk of small populations. However, it is generally unclear whether extinction risk from genetic causes gradually increases with decreasing population size or whether there is a sharp transition around a specific threshold population size. In the ecological literature, such threshold phenomena are called \"strong Allee effects\" and they can arise for example from mate limitation in small populations.\nC_LIO_LIIn this study, we aim to a) develop a meaningful notion of a \"strong genetic Allee effect\", b) explore whether and under what conditions such an effect can arise from inbreeding depression due to recessive deleterious mutations, and c) quantify the interaction of potential genetic Allee effects with the well-known mate-finding Allee effect.\nC_LIO_LIWe define a strong genetic Allee effect as a genetic process that causes a populations survival probability to be a sigmoid function of its initial size. The inflection point of this function defines the critical population size. To characterize survival-probability curves, we develop and analyze simple stochastic models for the ecology and genetics of small populations.\nC_LIO_LIOur results indicate that inbreeding depression can indeed cause a strong genetic Allee effect, but only if individuals carry sufficiently many deleterious mutations (lethal equivalents) on average and if these mutations are spread across sufficiently many loci. Populations suffering from a genetic Allee effect often first grow, then decline as inbreeding depression sets in, and then potentially recover as deleterious mutations are purged. Critical population sizes of ecological and genetic Allee effects appear to be often additive, but even superadditive interactions are possible.\nC_LIO_LIMany published estimates for the number of lethal equivalents in birds and mammals fall in the parameter range where strong genetic Allee effects are expected. Unfortunately, extinction risk due to genetic Allee effects can easily be underestimated as populations with genetic problems often grow initially, but then crash later. Also interactions between ecological and genetic Allee effects can be strong and should not be neglected when assessing the viability of endangered or introduced populations.\nC_LI

Ecology

Meta-GWAS Accuracy and Power (MetaGAP) calculator shows that hiding heritability is partially due to imperfect genetic correlations across studies

Large-scale genome-wide association results are typically obtained from a fixed-effects meta-analysis of GWAS summary statistics from multiple studies spanning different regions and/or time periods. This approach averages the estimated effects of genetic variants across studies. In case genetic effects are heterogeneous across studies, the statistical power of a GWAS and the predictive accuracy of polygenic scores are attenuated, contributing to the so-called missing heritability. Here, we describe the online Meta-GWAS Accuracy and Power calculator (MetaGAP; available at www.devlaming.eu) which quantifies this attenuation based on a novel multi-study framework. By means of simulation studies, we show that under a wide range of genetic architectures, the statistical power and predictive accuracy provided by this calculator are accurate. We compare the predictions from MetaGAP with actual results obtained in the GWAS literature. Specifically, we use genomic-relatedness-matrix restricted maximum likelihood (GREML) to estimate the SNP heritability and cross-study genetic correlation of height, BMI, years of education, and self-rated health in three large samples. These estimates are used as input parameters for the MetaGAP calculator. Results from the calculator suggest that cross-study heterogeneity has led to attenuation of statistical power and predictive accuracy in recent large-scale GWAS efforts on these traits (e.g., for years of education, we estimate a relative loss of 51-62% in the number of genome-wide significant loci and a relative loss in polygenic score R2 of 36-38%). Hence, cross-study heterogeneity contributes to the missing heritability.\n\nAuthor SummaryLarge-scale genome-wide association studies are uncovering the genetic architecture of traits which are affected by many genetic variants. Such studies typically meta-analyze association results from multiple studies spanning different regions and/or time periods. GWAS results do not yet capture a large share of the total proportion of trait variation attributable to genetic variation. The origins of this so-called missing heritability have been strongly debated. One factor exacerbating the missing heritability is heterogeneity in the effects of genetic variants across studies. Its influence on statistical power to detect associated genetic variants and the accuracy of polygenic predictions is poorly understood. In the current study, we derive the precise effects of heterogeneity in genetic effects across studies on both the statistical power to detect associated genetic variants as well as the accuracy of polygenic predictions. We provide an online calculator, available at www.devlaming.eu, which accounts for these effects. By means of this calculator, we show that imperfect genetic correlations between studies substantially decrease statistical power and predictive accuracy and, thereby, contribute to the missing heritability. The MetaGAP calculator helps researchers to gauge how sensitive their results will be to heterogeneity in genetic effects across studies. If strong heterogeneity is expected, random-instead of fixed-effects meta-analysis methods should be used.

Genetics

Shared genetics and couple-associated environment are major contributors to the risk of both clinical and self-declared depression

BackgroundBoth genetic and environmental contributions to risk of depression have been identified, but estimates of their effects are limited. Commonalities between major depressive disorder (MDD) and self-declared depression (SDD) are also unclear. Dissecting the genetic and environmental contributions to these traits and their correlation would inform the design and interpretation of genetic studies.\n\nMethodsUsing data from a large Scottish family-based cohort (GS:SFHS, N=21,387), we estimated the genetic and environmental contributions to MDD and SDD. Genetic effects associated with common genome-wide genetic variants (SNP heritability) and additional pedigree-associated genetic variation and Non-genetic effects associated with common environments were estimated using linear mixed modeling (LMM).\n\nFindingsBoth MDD and SDD had significant contributions from effects of common genetic variants, the additional genetic effect of the pedigree and the common environmental effect shared by couples. The correlation between SDD and MDD was high (r=1*00, se=0*21) for common-variant-associated genetic effects and moderate for both the additional genetic effect of the pedigree (r=0*58, se=0*08) and the couple-shared environmental effect (r=0*53, se=0*22).\n\nInterpretationBoth genetics and couple-shared environmental effects were the major factors influencing liability to depression. SDD may provide a scalable alternative to MDD in studies seeking to identify common risk variants. Rarer variants and environmental effects may however differ substantially according to different definitions of depression.\n\nFundingStudy supported by Wellcome Trust Strategic Award 104036/Z/14/Z. GS:SFHS funded by the Scottish Government Health Department, Chief Scientist Office, number CZD/16/6.

Genetics

The Genetic Architecture of Quantitative Traits Cannot Be Inferred From Variance Component Analysis

Classical quantitative genetic analyses estimate additive and non-additive genetic and environmental components of variance from phenotypes of related individuals. The genetic variance components are defined in terms of genotypic values reflecting underlying genetic architecture (additive, dominance and epistatic genotypic effects) and allele frequencies. However, the dependency of the definition of genetic variance components on the underlying genetic models is not often appreciated. Here, we show how the partitioning of additive and non-additive genetic variation is affected by the genetic models and parameterization of allelic effects. We show that arbitrarily defined variance components often capture a substantial fraction of total genetic variation regardless of the underlying genetic architecture in simulated and real data. Therefore, variance component analysis cannot be used to infer genetic architecture of quantitative traits. The genetic basis of quantitative trait variation in a natural population can only be defined empirically using high resolution mapping methods followed by detailed characterization of QTL effects.

Genetics

Genomic analysis of family data reveals additional genetic effects on intelligence and personality

Pedigree-based analyses of intelligence have reported that genetic differences account for 50-80% of the phenotypic variation. For personality traits these effects are smaller, with 34-48% of the variance being explained by genetic differences. However, molecular genetic studies using unrelated individuals typically report a heritability estimate of around 30% for intelligence and between 0% and 15% for personality variables. Pedigree-based estimates and molecular genetic estimates may differ because current genotyping platforms are poor at tagging causal variants, variants with low minor allele frequency, copy number variants, and structural variants. Using [~]20 000 individuals in the Generation Scotland family cohort genotyped for [~]700 000 single nucleotide polymorphisms (SNPs), we exploit the high levels of linkage disequilibrium (LD) found in members of the same family to quantify the total effect of genetic variants that are not tagged in GWASs of unrelated individuals. In our models, genetic variants in low LD with genotyped SNPs explain over half of the genetic variance in intelligence, education, and neuroticism. By capturing these additional genetic effects our models closely approximate the heritability estimates from twin studies for intelligence and education, but not for neuroticism and extraversion. We then replicated our finding using imputed molecular genetic data from unrelated individuals to show that [~]50% of differences in intelligence, and [~]40% of the differences in education, can be explained by genetic effects when a larger number of rare SNPs are included. From an evolutionary genetic perspective, a substantial contribution of rare genetic variants to individual differences in intelligence and education is consistent with mutation-selection balance.

genetics

Modeling of a negative feedback mechanism explains age-dependent genetic architecture in reproduction in domesticated C. elegans strains

Most biological traits and common diseases have a strong but complex genetic basis, controlled by large numbers of genetic variants with small contributions to a trait or disease risk. The effect-size of most genetic variants is not absolute, but can depend on a number of factors including the age and genetic background of an organism. In order to understand the mechanisms that cause these changes, we are studying heritable trait differences between two domesticated strains of C. elegans. We previously identified a major effect locus, caused by a mutation in a component of the NURF chromatin remodeling complex, that regulated reproductive output in an age-dependent manner. The effect-size of this locus changes from positive to negative over the course of an animals reproductive lifespan. Using a previously published macroscale model of egg-laying rate in C. elegans, we show how time-dependent effect-size can be explained by an unequal use of sperm combined with negative feedback between sperm and ovulation rate. We validate a number of key predictions of this model using controlled mating experiments and quantification of oogenesis and sperm use. By incorporating this model into QTL mapping, we identify and partition new QTLs into specific aspects of the egg-laying process. Finally, we show how epistasis between two genetic variants is predicted by this modeling as a consequence of unequal use of sperm. This work demonstrates how modeling of multicellular communication systems can improve our ability to predict and understand the role of genetic variation on a complex phenotype. Negative autoregulatory feedback loops, common in transcriptional regulation, could play an important role in modifying genetic architecture in other traits.\n\nAUTHOR SUMMARYComplex traits are influenced not only by the individual effects of genetic variants, but also how these variants interact with the environment, age, and each other. While complex genetic architectures seem to be ubiquitous in natural traits, little is known about the mechanisms that cause them. Here we identify an example of age-dependent genetic architecture controlling the rate and timing of reproduction in the hermaphroditic nematode C. elegans. Using computational modeling, we demonstrate how this age-dependent genetic architecture can arise as a consequence of two factors: hormonal feedback on oocytes mediated by major sperm protein (MSP) released by sperm stored in the spermatheca and life history differences in sperm use caused by genetic variants. Our work also suggests how age-dependent epistasis can emerge from multicellular feedback systems.

genetics

Inbred or Outbred? Genetic diversity in laboratory rodent colonies

Non-model rodents are widely used as subjects for both basic and applied biological research, but the genetic diversity of the study individuals is rarely quantified. University-housed colonies tend to be small and subject to founder effects and genetic drift and so may be highly inbred or show substantial genetic divergence from other colonies, even those derived from the same source. Disregard for the levels of genetic diversity in an animal colony may result in a failure to replicate results if a different colony is used to repeat an experiment, as different colonies may have fixed alternative variants. Here we use high throughput sequencing to demonstrate genetic divergence in three isolated colonies of Mongolian gerbil (Meriones unguiculatus) even though they were all established recently from the same source. We also show that genetic diversity in allegedly outbred colonies of non-model rodents (gerbils, hamsters, house mice, and deer mice) varies considerably from nearly no segregating diversity, to very high levels of polymorphism. We conclude that genetic divergence in isolated colonies may play an important role in the replication crisis. In a more positive light, divergent rodent colonies represent an opportunity to leverage genetically distinct individuals in genetic crossing experiments. In sum, awareness of the genetic diversity of an animal colony is paramount as it allows researchers to properly replicate experiments and also to capitalize on other, genetically distinct individuals to explore the genetic basis of a trait.

genetics

Optimal cross selection for long-term genetic gain in two-part programs with rapid recurrent genomic selection

This study evaluates optimal cross selection for balancing selection and maintenance of genetic diversity in two-part plant breeding programs with rapid recurrent genomic selection. The two-part program reorganizes a conventional breeding program into population improvement component with recurrent genomic selection to increase the mean of germplasm and product development component with standard methods to develop new lines. Rapid recurrent genomic selection has a large potential, but is challenging due to genotyping costs or genetic drift. Here we simulate a wheat breeding program for 20 years and compare optimal cross selection against truncation selection in the population improvement with one to six cycles per year. With truncation selection we crossed a small or a large number of parents. With optimal cross selection we jointly optimised selection, maintenance of genetic diversity, and cross allocation with AlphaMate program. The results show that the two-part program with optimal cross selection delivered the largest genetic gain that increased with the increasing number of cycles. With four cycles per year optimal cross selection had 78% (15%) higher long-term genetic gain than truncation selection with a small (large) number of parents. Higher genetic gain was achieved through higher efficiency of converting genetic diversity into genetic gain; optimal cross selection quadrupled (doubled) efficiency of truncation selection with a small (large) number of parents. Optimal cross selection also reduced the drop of genomic selection accuracy due to the drift between training and prediction populations. In conclusion, optimal cross-selection enables optimal management and exploitation of population improvement germplasm in two-part programs.\n\nKey messageOptimal cross selection increases long-term genetic gain of two-part programs with rapid recurrent genomic selection. It achieves this by optimising efficiency of converting genetic diversity into genetic gain through reducing the loss of genetic diversity and reducing the drop of genomic prediction accuracy with rapid cycling.

genetics

Genomic heritability estimates in sweet cherry reveal non-additive genetic variance is relevant for industry-prioritized traits

BackgroundSweet cherry is consumed widely across the world and provides substantial economic benefits in regions where it is grown. While cherry breeding has been conducted in the Pacific Northwest for over half a century, little is known about the genetic architecture of important traits. We used a genome-enabled mixed model to predict the genetic performance of 505 individuals for 32 phenological, disease response and fruit quality traits evaluated in the RosBREED sweet cherry crop data set. Genome-wide predictions were estimated using a repeated measures model for phenotypic data across 3 years, incorporating additive, dominance and epistatic variance components. Genomic relationship matrices were constructed with high-density SNP data and were used to estimate relatedness and account for incomplete replication across years.\n\nResultsHigh broad-sense heritabilities of 0.83, 0.77, and 0.75 were observed for days to maturity, firmness, and fruit weight, respectively. Epistatic variance exceeded 40% of the total genetic variance for maturing timing, firmness and powdery mildew response. Dominance variance was the largest for fruit weight and fruit size at 34% and 27%, respectively. Omission of non-additive sources of genetic variance from the genetic mode resulted in inflation of narrow-sense heritability but minimally influenced prediction accuracy of genetic values in validation. Predicted genetic rankings of individuals from single-year models were inconsistent across years, likely due to incomplete sampling of the population genetic variance.\n\nConclusionsPredicted breeding values and genetic values a measure revealed many high-performing individuals for use as parents and the most promising selections to advance for cultivar release consideration, respectively. This study highlights the importance of using the appropriate genetic model for calculating breeding values to avoid inflation of expected parental contribution to genetic gain. The genomic predictions obtained will enable breeders to efficiently leverage the genetic potential of North American sweet cherry germplasm by identifying high quality individuals more rapidly than with phenotypic data alone.

genetics

Genetic variability and potential effects on clinical trial outcomes: perspectives in Parkinson’s disease

BackgroundImproper randomization in clinical trials can result in the failure of the trial to meet its primary end-point. The last [~]10 years have revealed that common and rare genetic variants are an important disease factor and sometimes account for a substantial portion of disease risk variance. However, the burden of common genetic risk variants is not often considered in the randomization of clinical trials and can therefore lead to additional unwanted variance between trial arms. We simulated clinical trials to estimate false negative and false positive rates and investigated differences in single variants and mean genetic risk scores (GRS) between trial arms to investigate the potential effect of genetic variance on clinical trial outcomes at different sample sizes.\n\nMethodsSingle variant and genetic risk score analyses were conducted in a clinical trial simulation environment using data from 5851 Parkinsons Disease patients as well as two simulated virtual cohorts based on public data. The virtual cohorts included a GBA variant cohort and a two variant interaction cohort. Data was resampled at different sizes (n = 200-5000 for the Parkinsons Disease cohort) and (n = 50-800 and n = 50-2000 for virtual cohorts) for 1000 iterations and randomly assigned to the two arms of a trial. False negative and false positive rates were estimated using simulated clinical trials, and percent difference in genetic risk score and allele frequency was calculated to quantify disparity between arms.\n\nFindingsSignificant genetic differences between the two arms of a trial are found at all sample sizes. Approximately 90% of the iterations had at least one statistically significant difference in individual risk SNPs between each trial arm. Approximately 10% of iterations had a statistically significant difference between trial arms in polygenic risk score mean or variance. For significant iterations at sample size 200, the average percent difference for mean GRS between trial arms was 130.87%, decreasing to 29.87% as sample size reached 5000. In the GBA only simulations we see an average 18.86% difference in GRS scores between trial arms at n = 50, decreasing to 3.09% as sample size reaches 2000. Balancing patients by genotype reduced mean percent difference in GRS between arms to 36.71% for the main cohort and 2.00% for the GBA cohort at n = 200. When adding a drug effect to the simulations, we found that unbalanced genetics with an effect on the chosen measurable clinical outcome can result in high false negative rates among trials, especially at small sample sizes. At a sample size of n = 50 and a targeted drug effect of -0.5 points in UPDRS per year, we discovered 33.9% of trials resulted in false negatives.\n\nInterpretationsOur data support the hypothesis that within genetically unmatched clinical trials, particularly those below 1000 participants, heterogeneity could confound true therapeutic effects as expected. This is particularly important in the changing environment of drug approvals. Clinical trials should undergo pre-trial genetic adjustment or, at the minimum, post-trial adjustment and analysis for failed trials. Clinical trial arms should be balanced on genetic risk variants, as well as cumulative variant distributions represented by GRS, in order to ensure the maximum reduction in trial arm disparities. The reduction in variance after balancing allows smaller sample sizes to be utilized without risking the large disparities between trial arms witnessed in typical randomized trials. As the cost of genotyping will likely be far less than greatly increasing sample size, genetically balancing trial arms can lead to more cost-effective clinical trials as well as better outcomes.

genomics

Genetically increased serum calcium levels reduce Alzheimer’s disease risk

IMPORTANCE Alzheimers disease (AD) is the leading cause of disability in the elderly. It has been a long time about the calcium hypothesis of AD on the basis of emerging evidence since 1994. However, most studies focused on the association between calcium homeostasis and AD, and concerned the intracellular calcium concentration. Only few studies reported reduced serum calcium levels in AD. Until now, it remains unclear whether serum calcium levels are genetically associated with AD risk.\n\nOBJECTIVE To evaluate the genetic association between increased serum calcium levels and AD risk\n\nDESIGN, SETTING, AND PARTICIPANTS We performed a Mendelian randomization study to investigate the association of increased serum calcium with AD risk using the genetic variants from the large-scale serum calcium genome-wide association study (GWAS) dataset (N=61,079 individuals of European descent) and the large-scale AD GWAS dataset (N=54,162 individuals including 17,008 AD cases and 37,154 controls of European descent). Inverse-variance weighted meta-analysis (IVW) was used to provide a combined estimate of the genetic association. Meanwhile, we selected the weighted median regression and MR-Egger regression as the complementary analysis methods to examine the robustness of the IVW estimate.\n\nEXPOSURES Genetic predisposition to increased serum calcium levels\n\nMAIN OUTCOMES AND MEASURES The risk of AD.\n\nRESULTS We selected 6 independent genetic variants influencing serum calcium levels as the instrumental variables. IVW analysis showed that a genetically increased serum calcium level (per 1 standard deviation (SD) increase 0.5-mg/dL) was significantly associated with a reduced AD risk (OR=0.56, 95% CI: 0.34-0.94, P=5.00E-03). Meanwhile, both the weighted median estimate (OR=0.60, 95% CI: 0.34-1.06, P=0.08) and MR-Egger estimate (OR=0.66, 95% CI: 0.26-1.67, P=0.381) were consistent with the IVW estimate in terms of direction and magnitude.\n\nCONCLUSIONS AND RELEVANCE We provided evidence that genetically increased serum calcium levels could reduce the risk of AD. Meanwhile, randomized controlled study should be further conducted to assess the effect of serum calcium levels on AD risk, and further clarify whether diet calcium intake or calcium supplement, or both could reduce the risk of AD.\n\nKey PointsQuestion Is there a genetic relationship between elevated serum calcium levels and the risk of Alzheimers disease?\n\nFindings This Mendelian randomization study showed that the genetically increased serum calcium levels were associated with the reduced risk of Alzheimers disease.\n\nMeaning These findings provide evidence that genetically increased serum calcium levels could reduce the risk of Alzheimers disease.

genetics

Clustering of Type 2 Diabetes Genetic Loci by Multi-Trait Associations Identifies Disease Mechanisms and Subtypes

BackgroundType 2 diabetes (T2D) is a heterogeneous disease for which 1) disease-causing pathways are incompletely understood and 2) sub-classification may improve patient management. Unlike other biomarkers, germline genetic markers do not change with disease progression or treatment. In this paper we test whether a germline genetic approach informed by physiology can be used to deconstruct T2D heterogeneity. First, we aimed to categorize genetic loci into groups representing likely disease mechanistic pathways. Second, we asked whether the novel clusters of genetic loci we identified have any broad clinical consequence, as assessed in four independent cohorts of individuals with T2D.\n\nMethods and FindingsIn an effort to identify mechanistic pathways driven by established T2D genetic loci, we applied Bayesian nonnegative matrix factorization clustering to genome-wide association results for 94 independent T2D genetic loci and 47 diabetes-related traits. We identified five robust clusters of T2D loci and traits, each with distinct tissue-specific enhancer enrichment based on analysis of epigenomic data from 28 cell types. Two clusters contained variant-trait associations indicative of reduced beta-cell function, differing from each other by high vs. low proinsulin levels. The three other clusters displayed features of insulin resistance: obesity-mediated (high BMI, waist circumference), \"lipodystrophy-like\" fat distribution (low BMI, adiponectin, HDL-cholesterol, and high triglycerides), and disrupted liver lipid metabolism (low triglycerides). Increased cluster GRSs were associated with distinct clinical outcomes, including increased blood pressure, coronary artery disease, and stroke risk. We evaluated the potential for clinical impact of these clusters in four studies containing participants with T2D (METSIM, N=487; Ashkenazi, N=509; Partners Biobank, N=2,065; UK Biobank N=14,813). Individuals with T2D in the top genetic risk score decile for each cluster reproducibly exhibited the predicted cluster-associated phenotypes, with ~30% of all participants assigned to just one cluster top decile.\n\nConclusionOur approach identifies salient T2D genetically anchored and physiologically informed pathways, and supports use of genetics to deconstruct T2D heterogeneity. Classification of patients by these genetic pathways may offer a step toward genetically informed T2D patient management.

genetics

Genetic architecture and selective sweeps after polygenic adaptation to distant trait optima

Understanding the genetic basis of phenotypic adaptation to changing environments is an essential goal of population and quantitative genetics. While technological advances now allow interrogation of genome-wide genotyping data in large panels, our understanding of the process of polygenic adaptation is still limited. To address this limitation, we use extensive forward-time simulation to explore the impacts of variation in demography, trait genetics, and selection on the rate and mode of adaptation and the resulting genetic architecture. We simulate a population adapting to an optimum shift, modeling sequence variation for 20 QTL for each of 12 different demographies for 100 different traits varying in the effect size distribution of new mutations, the strength of stabilizing selection, and the contribution of the genomic background. We then use random forest regression approaches to learn the relative importance of input parameters in determining a number of aspects of the process of adaptation including the speed of adaptation, the relative frequency of hard sweeps and sweeps from standing variation, or the final genetic architecture of the trait. We find that selective sweeps occur even for traits under relatively weak selection and where the genetic background explains most of the variation. Though most sweeps occur from variation segregating in the ancestral population, new mutations can be important for traits under strong stabilizing selection that undergo a large optimum shift. We also show that population bottlenecks and expansion impact overall genetic variation as well as the relative importance of sweeps from standing variation and the speed with which adaptation can occur. We then compare our results to two traits under selection during maize domestication, showing that our simulations qualitatively recapitulate differences between them. Overall, our results underscore the complex population genetics of individual loci in even relatively simple quantitative trait models, but provide a glimpse into the factors that drive this complexity and the potential of these approaches for understanding polygenic adaptation.\n\nAuthor summaryMany traits are controlled by a large number of genes, and environmental changes can lead to shifts in trait optima. How populations adapt to these shifts depends on a number of parameters including the genetic basis of the trait as well as population demography. We simulate a number of trait architectures and population histories to study the genetics of adaptation to distant trait optima. We find that selective sweeps occur even in traits under relatively weak selection and our machine learning analyses find that demography and the effect sizes of mutations have the largest influence on genetic variation after adaptation. Maize domestication is a well suited model for trait adaptation accompanied by demographic changes. We show how two example traits under a maize specific demography adapt to a distant optimum and demonstrate that polygenic adaptation is a well suited model for crop domestication even for traits with major effect loci.

evolutionary biology

Fast and general-purpose linear mixed models for genome-wide genetics

Linear mixed effect models are powerful tools used to account for population structure in genome-wide association studies (GWASs) and estimate the genetic architecture of complex traits. However, fully-specified models are computationally demanding and common simplifications often lead to reduced power or biased inference. We describe Grid-LMM (https://github.com/deruncie/GridLMM), an extendable algorithm for repeatedly fitting complex linear models that account for multiple sources of heterogeneity, such as additive and non-additive genetic variance, spatial heterogeneity, and genotype-environment interactions. Grid-LMM can compute approximate (yet highly accurate) frequentist test statistics or Bayesian posterior summaries at a genome-wide scale in a fraction of the time compared to existing general-purpose methods. We apply Grid-LMM to two types of quantitative genetic analyses. The first is focused on accounting for spatial variability and non-additive genetic variance while scanning for QTL; and the second aims to identify gene expression traits affected by non-additive genetic variation. In both cases, modeling multiple sources of heterogeneity leads to new discoveries.\n\nAuthor summaryThe goal of quantitative genetics is to characterize the relationship between genetic variation and variation in quantitative traits such as height, productivity, or disease susceptibility. A statistical method known as the linear mixed effect model has been critical to the development of quantitative genetics. First applied to animal breeding, this model now forms the basis of a wide-range of modern genomic analyses including genome-wide associations, polygenic modeling, and genomic prediction. The same model is also widely used in ecology, evolutionary genetics, social sciences, and many other fields. Mixed models are frequently multi-faceted, which is necessary for accurately modeling data that is generated from complex experimental designs. However, most genomic applications use only the simplest form of linear mixed methods because the computational demands for model fitting can be too great. We develop a flexible approach for fitting linear mixed models to genome scale data that greatly reduces their computational burden and provides flexibility for users to choose the best statistical paradigm for their data analysis. We demonstrate improved accuracy for genetic association tests, increased power to discover causal genetic variants, and the ability to provide accurate summaries of model uncertainty using both simulated and real data examples.

genomics

Gut microbiota composition explains more variance in the host cardiometabolic risk than genetic ancestry

BackgroundCardiometabolic affections greatly contribute to the global burden of disease. The susceptibility to these conditions associates with the ancestral genetic composition and gut microbiota. However, studies explicitly testing associations between genetic ancestry and gut microbes are rare. We examined whether the ancestral genetic composition was associated with gut microbiota, and split apart the effects of genetic and non-genetic factors on host health.\n\nResultsWe performed a cross-sectional study of 441 community-dwelling Colombian mestizos from five cities. We characterized the host genetic ancestry using 40 ancestry informative markers and gut microbiota through 16S rRNA gene sequencing. We measured variables related to cardiometabolic health (adiposity, blood chemistry and blood pressure), diet (calories, macronutrients and fiber) and lifestyle (physical activity, smoking and medicament consumption). The ancestral genetic composition of the studied population was 67{+/-}6% European, 21{+/-}5% Native American and 12{+/-}5% African. While we found limited evidence of associations between genetic ancestry and gut microbiota or disease risk, we observed a strong link between gut microbes and cardiometabolic health. Multivariable-adjusted linear models indicated that gut microbiota was more likely to explain variance in host health than genetic ancestry. Further, we identified 9 OTUs associated with increased disease risk and 11 with decreased risk.\n\nConclusionsGut microbiota seems to be more meaningful to explain cardiometabolic disease risk than genetic ancestry in this mestizo population. Our study suggests that novel ways to control cardiometabolic disease risk, through modulation of the gut microbial community, could be applied regardless of the genetic ancestry of the intervened population.

microbiology

Differential relationships between habitat fragmentation and within-population genetic diversity of three forest-dwelling birds

Habitat fragmentation is a major driver of environmental change affecting wildlife populations across multiple levels of biological diversity. Much of the recent research in landscape genetics has focused on quantifying the influence of fragmentation on genetic variation among populations, but questions remain as to how habitat loss and configuration influences within-population genetic diversity. Habitat loss and fragmentation might lead to decreases in genetic diversity within populations, which might have implications for population persistence over multiple generations. We used genetic data collected from populations of three species occupying forested landscapes across a broad geographic region: Mountain Chickadee (Poecile gambeli; 22 populations), White-breasted Nuthatch (Sitta carolinensis; 13 populations) and Pygmy Nuthatch (Sitta pygmaea; 19 populations) to quantify patterns of haplotype and nucleotide diversity across a range of forest fragmentation. We predicted that fragmentation effects on genetic diversity would vary depending on dispersal capabilities and habitat specificity of the species. Forest aggregation and the variability in forest patch area were the two strongest landscape predictors of genetic diversity. We found higher haplotype diversity in populations of P. gambeli and S. carolinensis inhabiting landscapes characterized by lower levels of forest fragmentation. Conversely, S. pygmaea demonstrated the opposite pattern of higher genetic diversity in fragmented landscapes. For two of the three species, we found support for the prediction that highly fragmented landscapes sustain genetically less diverse populations. We suggest, however, that future studies should focus on species of varying life-history traits inhabiting independent landscapes to better understand how habitat fragmentation influences within-population genetic diversity.

Genetics

Analysis of genetic similarity among friends and schoolmates in the National Longitudinal Study of Adolescent to Adult Health (Add Health)

Humans tend to form social relationships with others who resemble them. Whether this sorting of like with like arises from historical patterns of migration, meso-level social structures in modern society, or individual-level selection of similar peers remains unsettled. Recent research has evaluated the possibility that unobserved genotypes may play an important role in the creation of homophilous relationships. We extend this work by using data from 9,500 adolescents from the National Longitudinal Study of Adolescent to Adult Health (Add Health) to examine genetic similarities among pairs of friends. While there is some evidence that friends have correlated genotypes, both at the whole-genome level as well as at trait-associated loci (via polygenic scores), further analysis suggests that meso-level forces, such as school assignment, are a principal source of genetic similarity between friends. We also observe apparent social-genetic effects in which polygenic scores of an individuals friends and schoolmates predict the individuals own educational attainment. In contrast, an individuals height is unassociated with the height genetics of peers.\n\nSignificanceOur study reported significant findings of a \"social genome\" that can be quantified and studied to understand human health and behavior. In a national sample of more than 9,000 American adolescents, we found evidence of social forces that act to make friends and schoolmates more genetically similar to one another as compared to random pairs of unrelated individuals. This subtle genetic similarity was observed across the entire genome and at sets of genomic locations linked with specific traits--educational attainment and body-mass index--a phenomenon we term \"social-genetic correlation.\" We also find evidence of a \"social-genetic effect\" such that the genetics of a persons friends and schoolmates influenced their own education, even after accounting for the persons own genetics.

genetics

The genetics of the human face: identification of large effect single gene variants

In order to discover specific variants with relatively large effects on the human face we have devised an approach to identifying facial features with high heritability. This is based on using twin data to estimate the additive genetic value of each point on a face, as provided by a 3D camera system. In addition, we have used the ethnic difference between East Asian and European faces as a further source of face genetic variation. We use principal components analysis to provide a fine definition of the surface features of human faces around the eyes and of the profile, and chose upper and lower 10% extremes of the most heritable PCs for looking for genetic associations. Using this strategy for the analysis of 3D images of 1832 unique volunteers from the well characterised People of the British Isles study [1, 2] and 1567 unique twin images from the TwinsUK cohort (www.twinsuk.ac.uk), together with genetic data for 500,000 SNPs, we have identified three specific genetic variants with notable effects on facial profiles and eyes.\n\nSignificance statementThe human face is extraordinarily variable and the extreme similarity of the faces of identical twins indicates that most of this variability is genetically determined. This level of genetic variability has probably arisen through natural selection, for example, for recognition of membership of a group or as a consequence of differential mate selection with respect to facial features. We have devised an approach to identifying specific genetic effects on particular facial features. This should enable the understanding, eventually at the molecular level, of the nature of this extraordinary genetic variability, which is such an important feature of our everyday human interactions.\n\nAuthor ContributionsWFB conceived the project. BW and TD organised collection of PoBI data and KH, DJMC, DM, AB and WFB assisted in data collection. TDS, PH and AN collected TwinsUK data. WPK and WJC conducted image registration analysis under supervision of JK. DJMC analysed registered image data and genetic data under supervision of WFB. WFB and DJMC wrote the manuscript, with WJC and WPK contributing additional technical material. WFB and BW supervised the project.

genetics