bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Quantitative approaches to variant classification increase the yield and precision of genetic testing in Mendelian diseases: The case of hypertrophic cardiomyopathy

BackgroundInternational guidelines for variant interpretation in Mendelian disease set stringent criteria to report a variant as (likely) pathogenic, prioritising control of false positive rate over test sensitivity and diagnostic yield. Genetic testing is also more likely informative in individuals with well-characterised variants from extensively studied European-ancestry populations. Inherited cardiomyopathies are relatively common Mendelian diseases that allow empirical calibration and assessment of this framework.\n\nResultsWe compared rare variants in large hypertrophic cardiomyopathy (HCM) cohorts to reference populations to identify variant classes with high prior likelihoods of pathogenicity, as defined by etiological fraction (EF). Analysis of variant distribution identified regions in which variants are significantly enriched in cases and variant location was a better discriminator of pathogenicity than generic computational functional prediction algorithms. Non-truncating variant classes with an EF[≥]0.95, and therefore clinically actionable, were identified in 5 established HCM genes. Applying this approach leads to an estimated 14-20% increase in cases with actionable HCM variants.\n\nConclusionsWhen found in a patient confirmed to have disease, novel variants in some genes and regions are empirically shown to have a sufficiently high probability of pathogenicity to support a \"likely pathogenic\" classification, even without additional segregation or functional data. This could increase the yield of high confidence actionable variants, consistent with the framework and recommendations of current guidelines. The techniques outlined offer a consistent, unbiased and equitable approach to variant interpretation for Mendelian disease genetic testing. We propose adaptations to ACMG/AMP guidelines to incorporate such evidence in a quantitative and transparent manner.

genetics

Measuring intolerance to mutation in human genetics

In numerous applications, from working with animal models to mapping the genetic basis of human disease susceptibility, it is useful to know whether a single disrupting mutation in a gene is likely to be deleterious1-4. With this goal in mind, a number of measures have been developed to identify genes in which protein-truncating variants (PTVs), or other types of mutations, are absent or kept at very low frequency in large population samples--genes that appear \"intolerant to mutation\"3,5-9. One measure in particular, pLI, has been widely adopted7. By contrasting the observed versus expected number of PTVs, it aims to classify genes into three categories, labelled \"null\", \"recessive\" and \"haploinsufficient\"7. Such population genetic approaches can be useful in many applications. As we clarify, however, these measures reflect the strength of selection acting on heterozygotes, and not dominance for fitness or haploinsufficiency for other phenotypes.

genetics

Dissection of complex, fitness-related traits in multiple Drosophila mapping populations offers insight into the genetic control of stress resistance

We leverage two complementary Drosophila melanogaster mapping panels to genetically dissect starvation resistance, an important fitness trait. Using >1600 genotypes of the multiparental Drosophila Synthetic Population Resource (DSPR) we map numerous starvation stress QTL that collectively explain a substantial fraction of trait heritability. QTL effects further allowed us to estimate DSPR founder phenotypes, predictions that were correlated with the actual founder phenotypes. Starvation resistance has been linked to triglyceride level, and while we observe a modest phenotypic correlation between the traits in the DSPR, overlap among the QTL identified for each trait is low. Since we show that DSPR strains with extreme starvation phenotypes also differ in desiccation resistance and activity level, our data imply that multiple physiological mechanisms contribute to starvation variability. We also exploit the Drosophila Genetic Reference Panel (DGRP), identifying a number of sequence variants associated with starvation resistance. Consistent with prior work these sites rarely fall within QTL intervals mapped in the DSPR. Two other groups previously measured starvation resistance in the DGRP, offering a unique opportunity to directly compare mapping results across labs. We found strong phenotypic correlations among studies, but extremely low overlap in the sets of genomewide significant sites. Despite this, our analyses revealed that the most highly-associated variants from each study typically showed the same additive effect sign in independent studies, in contrast to otherwise equivalent sets of random variants. This consistency provides evidence for reproducible trait-associated sites in a widely-used mapping panel, and highlights the polygenic nature of starvation resistance.

genetics

Efficient implementation of penalized regression for genetic risk prediction

Polygenic Risk Scores (PRS) consist in combining the information across many single-nucleotide polymorphisms (SNPs) in a score reflecting the genetic risk of developing a disease. PRS might have a major impact on public health, possibly allowing for screening campaigns to identify high-genetic risk individuals for a given disease. The \"Clumping+Thresholding\" (C+T) approach is the most common method to derive PRS. C+T uses only univariate genome-wide association studies (GWAS) summary statistics, which makes it fast and easy to use. However, previous work showed that jointly estimating SNP effects for computing PRS has the potential to significantly improve the predictive performance of PRS as compared to C+T.\n\nIn this paper, we present an efficient method to jointly estimate SNP effects, allowing for practical application of penalized logistic regression (PLR) on modern datasets including hundreds of thousands of individuals. Moreover, our implementation of PLR directly includes automatic choices for hyper-parameters. The choice of hyper-parameters for a predictive model is very important since it can dramatically impact its predictive performance. As an example, AUC values range from less than 60% to 90% in a model with 30 causal SNPs, depending on the p-value threshold in C+T.\n\nWe compare the performance of PLR, C+T and a derivation of random forests using both real and simulated data. PLR consistently achieves higher predictive performance than the two other methods while being as fast as C+T. We find that improvement in predictive performance is more pronounced when there are few effects located in nearby genomic regions with correlated SNPs; for instance, AUC values increase from 83% with the best prediction of C+T to 92.5% with PLR. We confirm these results in a data analysis of a case-control study for celiac disease where PLR and the standard C+T method achieve AUC of 89% and of 82.5%.\n\nIn conclusion, our study demonstrates that penalized logistic regression can achieve more discriminative polygenic risk scores, while being applicable to large-scale individual-level data thanks to the implementation we provide in the R package bigstatsr.

genetics

Estimation of realized rates of genetic gain and indicators for breeding program assessment

Routine estimation of the rate of genetic gain ({Delta}Gt) realized by a breeding program has been proposed as a means to monitor its effectiveness. Several methods of realized{Delta} Gt estimation have been utilized in other studies, but none have been objectively evaluated in a plant breeding context. Stochastic simulations of 80 rice (Oryza sativa) breeding programs over 28 years were done to generate data used to evaluate five methods of realized{Delta} Gt estimation in terms of error, precision, efficiency and correlation between true and predicted annual mean breeding values. Two indicators of{Delta} Gt, the expected{Delta} Gt and the average number of equivalent complete generations (EqCg), were described and evaluated. At best, estimates of realized{Delta} Gt were over or underestimated by 15% and 27% when considering all 28 years and the past 15 years of breeding respectively. The best methods were the control population, estimated breeding value, and ERA trial methods. Among these, correlations between true and estimated{Delta} Gt were at best 0.59, indicating that these methods cannot very accurately rank breeding programs in terms of realized{Delta} Gt. The expected{Delta} Gt and the average EqCg were shown to be useful indicators for determining if a non-zero genetic gain is expected. Determining which of the three best realized{Delta} Gt estimation methods evaluated, if any, would be appropriate for any given breeding program should be done with careful consideration of the objectives, resources, seed stocks, and structure of the data available.

genetics

Exploring Genetic Variation That Influences Brain Methylation In Attention-Deficit/Hyperactivity Disorder

Attention-deficit/hyperactivity disorder (ADHD) is a neurodevelopmental disorder caused by an interplay of genetic and environmental factors. Epigenetics is crucial to lasting changes in gene expression in the brain. Recent studies suggest a role for DNA methylation in ADHD. We explored the contribution to ADHD of allele-specific methylation (ASM), an epigenetic mechanism that involves SNPs correlating with differential levels of DNA methylation at CpG sites. We selected 3,896 tagSNPs reported to influence methylation in human brain regions and performed a case-control association study using the summary statistics from the largest GWAS meta-analysis of ADHD, comprising 20,183 cases and 35,191 controls. We identified associations with eight tagSNPs that were significant at a 5% False Discovery Rate (FDR). These SNPs correlated with methylation of CpG sites lying in the promoter regions of six genes. Since methylation may affect gene expression, we inspected these ASM SNPs together with 52 ASM SNPs in high LD with them for eQTLs in brain tissues and observed that the expression of three of those genes was affected by them. ADHD risk alleles correlated with increased expression (and decreased methylation) of ARTN and PIDD1 and with a decreased expression (and increased methylation) of C2orf82. Furthermore, these three genes were predicted to have altered expression in ADHD, and genetic variants in C2orf82 correlated with brain volumes. In summary, we followed a systematic approach to identify risk variants for ADHD that correlated with differential cis-methylation, identifying three novel genes contributing to the disorder.

genetics

Quantifying Heterogeneity in the Genetic Architecture of Complex Traits Between Ethnically Diverse Groups using Random Effect Interaction Models

In humans, most genome-wide association studies have been conducted using data from Caucasians and many of the reported findings have not replicated in other populations. This lack of replication may be due to statistical issues (small sample size, confounding) or perhaps more fundamentally to differences in the genetic architecture of traits between ethnically diverse subpopulations. What aspects of the genetic architecture of traits vary between subpopulations and how can this be quantified? We consider studying effect heterogeneity using random-effect Bayesian interaction models. The proposed methodology can be applied using shrinkage and variable selection methods and produces useful information about effect heterogeneity in the form of whole-genome summaries (e.g., SNP-heritability and the average correlation of effects) as well as SNP-specific attributes. Using simulations, we show that the proposed methodology yields (nearly) unbiased estimates of genomic heritability and of the average correlation of effects between groups when the sample size is not too small relative to the number of SNPs used. Subsequently, we used the proposed methodology for the analyses of four complex human traits (standing height, high-density lipoprotein, low-density lipoprotein, and serum urate levels) in European-Americans (EAs) and African-Americans (AAs). The estimated correlations of effects between the two subpopulations was well below unity for all the traits, ranging from 0.73 to 0.50. The extent of effect heterogeneity varied between traits and SNP-sets. Height showed less differences in SNP effects between AAs and EAs whereas HDL, a trait highly influenced by life-style, exhibited greater extent of effect heterogeneity. For all the traits we observed substantial variability in effect heterogeneity across SNPs, suggesting it varies between regions of the genome.

genetics

The genetic legacy of continental scale admixture in Indian Austroasiatic speakers

Surrounded by speakers of Indo-European, Dravidian and Tibeto-Burman languages, around 11 million Munda (a branch of Austroasiatic language family) speakers live in the densely populated and genetically diverse South Asia. Their genetic makeup holds components characteristic of South Asians as well as Southeast Asians. The admixture time between these components has been previously estimated on the basis of archaeology, linguistics and uniparental markers. Using genome-wide genotype data of 102 Munda speakers and contextual data from South and Southeast Asia, we retrieved admixture dates between 2000 - 3800 years ago for different populations of Munda. The best modern proxies for the source populations for the admixture with proportions 0.78/0.22 are Lao people from Laos and Dravidian speakers from Kerala in India, while the South Asian population(s), with whom the incoming Southeast Asians intermixed, had a smaller proportion of West Eurasian component than contemporary proxies. Somewhat surprisingly Malaysian Peninsular tribes rather than the geographically closer Austroasiatic languages speakers like Vietnamese and Cambodians show highest sharing of IBD segments with the Munda. In addition, we affirmed that the grouping of the Munda speakers into North and South Munda based on linguistics is in concordance with genome-wide data.

genetics

Survival protection mechanisms and genetic variability induction after stress: two sides of the same Hsp70 coin

Previous studies have shown that heat shock stress may increase transcription levels and, in some cases, also the transposition of certain transposable elements (TEs) in Drosophila and other organisms. Other studies have also demonstrated that heat shock chaperones as Hsp90 and Hop are involved in repressing transposons activity in Drosophila melanogaster by their involvement in crucial steps of the biogenesis of Piwi-interacting RNAs (piRNAs), the largest class of germline-enriched small non-coding RNA implicated in the epigenetic silencing of TEs. However, a satisfying picture of how many chaperones and their respective functional roles could be involved in repressing transposons in germ cell is still unknown. Here we show that in Drosophila heat shock activates transposon's expression at post-transcriptional level by disrupting a repressive chaperone complex by a decisive role of the stress-inducible chaperone Hsp70. We found that stress-induced transposons activation is triggered by an interaction of Hsp70 with the Hsc70-Hsp90 complex and other factors all involved in piRNA biogenesis in both ovaries and testes. Such interaction induces a displacement of all such factors to the lysosomes resulting in a functional collapse of piRNA biogenesis. In support of a significant role of Hsp70 in transposon activation after stress, we found that the expression under normal conditions of Hsp70 in transgenic flies increases the amount of transposon transcripts and displaces the components of chaperon machinery outside the nuage as observed after heat shock. So that, our results demonstrate that heat shock stress is capable to increase the expression of transposons at post-transcriptional level by affecting piRNA biogenesis through the action of the inducible chaperone Hsp70. We think that such mechanism proposes relevant evolutionary implications. In presence of drastic environmental changes, Hsp70 plays a key dual role in increasing both the survival probability of individuals and the genetic variability in their germ cells. This in turn should be translated into an increase of genetic variability inside the populations thus potentiating their evolutionary plasticity and evolvability.

genetics

Identifying novel subtypes of irritability using a developmental genetic approach

ObjectiveIrritability is a common reason for referral to services, strongly associated with impairment and negative outcomes, but is a nosological and treatment challenge. A major issue is how irritability should be conceptualized. This study used a developmental approach to test the hypothesis that there are several forms of irritability, including a neurodevelopmental/ADHD-like subtype with onset in childhood and a depression/mood subtype with onset in adolescence.\n\nMethodData were analyzed in the Avon Longitudinal Study of Parents and Children, a prospective UK population-based cohort. Irritability trajectory-classes were estimated for 7924 individuals with data at multiple time-points across childhood and adolescence (4 possible time-points from approximately ages 7 to 15 years). Psychiatric diagnoses were assessed at approximately ages 7 and 15 years. Psychiatric genetic risk was indexed by polygenic risk scores (PRS) for attention-deficit/hyperactivity disorder (ADHD) and major depressive disorder (MDD) derived using large genome-wide association study results.\n\nResultsFive irritability trajectory classes were identified: low (81.2%), decreasing (5.6%), increasing (5.5%), late-childhood limited (5.2%) and high-persistent (2.4%). The early-onset, high-persistent trajectory was associated with male preponderance, childhood ADHD (OR=108.64 (57.45-204.41), p<0.001) and ADHD PRS (OR=1.31 (1.09-1.58), p=0.005); the adolescent-onset, increasing trajectory was associated with female preponderance, adolescent MDD (OR=5.14 (2.47-10.73), p<0.001) and MDD PRS (OR=1.20, (1.05-1.38), p=0.009). Both trajectory classes were associated with MDD diagnosis and ADHD genetic risk.\n\nConclusionsThe developmental context of irritability may be important in its conceptualization: early-onset persistent irritability maybe more neurodevelopmental/ADHD-like and later-onset irritability more depression/mood-like. This has implications for treatment as well as nosology.

genetics

Applicability of the mutation-selection balance model to population genetics of heterozygous protein-truncating variants in humans

The fate of alleles in the human population is believed to be highly affected by the stochastic force of genetic drift. Estimation of the strength of natural selection in humans generally necessitates a careful modeling of drift including complex effects of the population history and structure. Protein truncating variants (PTVs) are expected to evolve under strong purifying selection and to have a relatively high per-gene mutation rate. Thus, it is appealing to model the population genetics of PTVs under a simple deterministic mutation-selection balance, as has been proposed earlier [1]. Here, we investigated the limits of this approximation using both computer simulations and data-driven approaches. Our simulations rely on a model of demographic history estimated from 33,370 individual exomes of the Non-Finnish European subset of the ExAC dataset [2]. Additionally, we compared the African and European subset of the ExAC study and analyzed de novo PTVs. We show that the mutation-selection balance model is applicable to the majority of human genes, but not to genes under the weakest selection.

genetics

The molecular genetics of hand preference revisited

Hand preference is a prominent behavioural trait linked to human brain asymmetry. A handful of genetic variants have been reported to associate with hand preference or quantitative measures related to it. Most of these reports were on the basis of limited sample sizes, by current standards for genetic analysis of complex traits. Here we performed a genome-wide association analysis of hand preference in the large, population-based UK Biobank cohort (N=331,037). We used gene-set enrichment analysis to investigate whether genes involved in visceral asymmetry are particularly relevant to hand preference, following one previous report. We found no evidence implicating any specific candidate variants previously reported. We also found no evidence that genes involved in visceral laterality play a role in hand preference. It remains possible that some of the previously reported genes or pathways are relevant to hand preference as assessed in other ways, or else are relevant within specific disorder populations. However, some or all of the earlier findings are likely to be false positives, and none of them appear relevant to hand preference as defined categorically in the general population. Within the UK Biobank itself, a significant association implicates the gene MAP2 in handedness.

genetics

Genetic Consequences of Social Stratification in Great Britain

Human DNA varies across geographic regions, with most variation observed so far reflecting distant ancestry differences. Here, we investigate the geographic clustering of genetic variants that influence complex traits and disease risk in a sample of ~450,000 individuals from Great Britain. Out of 30 traits analyzed, 16 show significant geographic clustering at the genetic level after controlling for ancestry, likely reflecting recent migration driven by socio-economic status (SES). Alleles associated with educational attainment (EA) show most clustering, with EA-decreasing alleles clustering in lower SES areas such as coal mining areas. Individuals that leave coal mining areas carry more EA-increasing alleles on average than the rest of Great Britain. In addition, we leveraged the geographic clustering of complex trait variation to further disentangle regional differences in socio-economic and cultural outcomes through genome-wide association studies on publicly available regional measures, namely coal mining, religiousness, 1970/2015 general election outcomes, and Brexit referendum results.

genetics

Largest genome-wide association study for PTSD identifies genetic risk loci in European and African ancestries and implicates novel biological pathways

Post-traumatic stress disorder (PTSD) is a common and debilitating disorder. The risk of PTSD following trauma is heritable, but robust common variants have yet to be identified by genome-wide association studies (GWAS). We have collected a multi-ethnic cohort including over 30,000 PTSD cases and 170,000 controls. We first demonstrate significant genetic correlations across 60 PTSD cohorts to evaluate the comparability of these phenotypically heterogeneous studies. In this largest GWAS meta-analysis of PTSD to date we identify a total of 6 genome-wide significant loci, 4 in European and 2 in African-ancestry analyses. Follow-up analyses incorporated local ancestry and sex-specific effects, and functional studies. Along with other novel genes, a non-coding RNA (ncRNA) and a Parkinsons Disease gene, PARK2, were associated with PTSD. Consistent with previous reports, SNP-based heritability estimates for PTSD range between 10-20%. Despite a significant shared liability between PTSD and major depressive disorder, we show evidence that some of our loci may be specific to PTSD. These results demonstrate the role of genetic variation contributing to the biology of differential risk for PTSD and the necessity of expanding GWAS beyond European ancestry.

genetics

AVADA Enables Automated Genetic Variant Curation Directly from the Full Text Literature

PurposeThe primary literature on human genetic diseases includes descriptions of pathogenic variants that are essential for clinical diagnosis. Variant databases such as ClinVar and HGMD collect pathogenic variants by manual curation. We aimed to automatically construct a freely accessible database of pathogenic variants directly from full-text articles about genetic disease.\n\nMethodsAVADA (Automatically curated VAriant DAtabase) is a novel machine learning tool that uses natural language processing to automatically identify pathogenic variants and genes in full text of primary literature and converts them to genomic coordinates for rapid downstream use.\n\nResultsAVADA automatically curated almost 60% of pathogenic variants deposited in HGMD, a 4.4-fold improvement over the current state of the art in automated variant extraction. AVADA also contains more than 60,000 pathogenic variants that are in HGMD, but not in ClinVar. In a cohort of 245 diagnosed patients, AVADA correctly annotated 38 previously described diagnostic variants, compared to 43 using HGMD, 20 using ClinVar and only 13 (wholly subsumed by AVADA and ClinVars) using the best automated abstracts-only based approach.\n\nConclusionAVADA is the first machine learning tool that automatically curates a variants database directly from full text literature. AVADA is available upon publication at http://bejerano.stanford.edu/AVADA.

genetics

On ‘Reverse’ Regression for Robust Genetic Association Studies and Allele Frequency Estimation with Related Individuals

For genetic association studies with related individuals, standard linear mixed-effect model is the most popular approach. The model treats a complex trait (phenotype) as the response variable while a genetic variant (genotype) as a covariate. An alternative approach is to reverse the roles of phenotype and genotype. This class of tests includes quasi-likelihood based score tests. In this work, after reviewing these existing methods, we propose a general, unifying reverse regression framework. We then show that the proposed method can also explicitly adjust for potential departure from Hardy-Weinberg equilibrium. Lastly, we demonstrate the additional flexibility of the proposed model on allele frequency estimation, as well as its connection with earlier work of best linear unbiased allele-frequency estimator. We conclude the paper with supporting evidence from simulation and application studies.

genetics

Subset selection of markers for genome-enabled prediction of genetic val-ues using radial basis function neural networks

This paper aimed to evaluate the efficiency of subset selection of markers for genome-enabled prediction of genetic values using radial basis function neural networks (RBFNN). For this purpose, an F1 population from hybridization of divergent parents with 500 individuals genotyped with 1,000 SNP-type markers was simulated. Phenotypic traits were determined by adopting three different gene action models - additive, additive-dominant, and epistasic, complying with two dominance situations: partial and complete with quantitative traits admitting heritability (h2) equal to 30 and 60%, each one controlled by 50 loci, considering two alleles per locus, totaling 12 different scenarios. To evaluate the predictive ability of RR_BLUP and the neural networks, a cross-validation procedure with five replicates were trained using 80% of the individuals of the population. Two methods were used: dimensionality reduction and stepwise regression. The square of the correlation between the predicted genomic estimated breeding value (GEBV) and the phenotype value was used to measure predictive reliability. For h2 = 0.3 in the additive scenario, the R2 values were 59% for neural network (RBFNN) and 57% for RR-BLUP, and in the epistatic scenario, R2 values were 50% and 41%, respectively. Additionally, when analyzing the mean-squared error root, the difference in performance between the techniques is even greater. For the additive scenario, the estimates were 91 for RR-BLUP and 5 for neural networks and, in the most critical scenario, they were 427 for RR-BLUP and 20 for neural network. The results showed that the use of neural networks and variable selection techniques allows capturing epistasis interactions, leading to an improvement in the accuracy of prediction of the genetic value and, mainly, to a large reduction of the mean square error, which indicates greater genomic value.

genetics

Genetics of fasting indices of glucose homeostasis using GWIS unravels tight relationships with inflammatory markers

PurposeHomeostasis Model Assessment of {beta}-cell function and Insulin Resistance (HOMA-B/-IR) indices are informative about the pathophysiological processes underlying type 2 diabetes (T2D). Data on both fasting glucose and insulin levels are required to calculate HOMA-B/-IR, leading to underpowered Genome-Wide Association studies (GWAS) of these traits.\n\nMethodsWe overcame such power loss issues by implementing Genome-Wide Inferred Statistics (GWIS) approach and subsequent dense genome-wide imputation of HOMA-B/-IR summary statistics with SS-imp to 1000 Genomes project variant density, reaching an analytical sample size of 75,240 European individuals without diabetes. We dissected mechanistic heterogeneity of glycaemic trait/T2D loci effects on HOMA-B/-IR and their relationships with 36 inflammatory and cardiometabolic phenotypes.\n\nResultsWe identified one/three novel HOMA-B (FOXA2)/HOMA-IR (LYPLAL1, PER4, PPP1R3B) loci. We detected novel strong genetic correlations between HOMA-IR/-B and Plasminogen Activator Inhibitor 1 (PAI-1, rg=0.92/0.78, P=2.13x10-4/2.54x10-3). HOMA-IR/-B were also correlated with C-Reactive Protein (rg=0.33/0.28, P=4.67x10-3/3.65x10-3). HOMA-IR was additionally correlated with T2D (rg=0.56, P=2.31x10-9), glycated haemoglobin (rg=0.28, P=0.024) and adiponectin (rg=-0.30, P=0.012).\n\nConclusionUsing innovative GWIS approach for composite phenotypes we report novel evidence for genetic relationships between fasting indices of insulin resistance/beta-cell function and inflammatory markers, providing further support for the role of inflammation in T2D pathogenesis.

genetics