bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Using imputed genotype data in the joint score tests for genetic association and gene-environment interactions in case-control studies

BackgroundGenome-wide association studies (GWAS) are now routinely imputed for untyped SNPs based on various powerful statistical algorithms for imputation trained on reference datasets. The use of predicted allele count for imputed SNPs as the dosage variable is known to produce valid score test for genetic association.\n\nMethodsIn this paper, we investigate how to best handle imputed SNPs in various modern complex tests for genetic association incorporating gene-environment interactions. We focus on case-control association studies where inference in an underlying logistic regression model can be performed using alternative methods that rely on varying degree on an assumption of gene-environment independence in the underlying population. As increasingly large scale GWAS are being performed through consortia effort where it is preferable to share only summary-level information across studies, we also describe simple mechanisms for implementing score-tests based on standard meta-analysis of \"one-step\" maximum-likelihood estimates across studies.\n\nResultsApplications of the methods in simulation studies and a dataset from genome-wide association study of lung cancer illustrate ability of the proposed methods to maintain type-I error rates for underlying testing procedures. For analysis of imputed SNPs, similar to typed SNPs, retrospective methods can lead to considerable efficiency gain for modeling of gene-environment interactions under the assumption of gene-environment independence.\n\nConclusionsProposed methods allow valid analysis of imputed SNPs in case-control studies of gene-environment interaction using alternative strategies that had been earlier available only for genotyped SNPs.

Genetics

High-resolution DNA accessibility profiles increase the discovery and interpretability of genetic associations

Genetic risk for common autoimmune diseases is influenced by hundreds of small effect, mostly non-coding variants, enriched in regulatory regions active in adaptive-immune cell types. DNaseI hypersensitivity sites (DHSs) are a genomic mark for regulatory DNA. Here, we generated a single DHSs annotation from fifteen deeply sequenced DNase-seq experiments in adaptive-immune as well as non-immune cell types. Using this annotation we quantified accessibility across cell types in a matrix format amenable to statistical analysis, deduced the subset of DHSs unique to adaptive-immune cell types, and grouped DHSs by cell-type accessibility profiles. Measuring enrichment with cell-type-specific TF binding sites as well as proximal gene expression and function, we show that accessibility profiles grouped DHSs into coherent regulatory functions. Using the adaptive-immune-specific DHSs as input (0.37% of genome), we associated DHSs to six autoimmune diseases with GWAS data. Associated loci showed higher replication rates when compared to loci identified by GWAS or by considering all DHSs, allowing the additional discovery of 327 loci (FDR<0.005) below typical GWAS significance threshold, 52 of which are novel and replicating discoveries. Finally, we integrated DHS associations from six autoimmune diseases, using a network model (bird-eye view) and a regulatory Manhattan plot schema (per locus). Taken together, we described and validated a strategy to leverage finely resolved regulatory priors, enhancing the discovery, interpretability, and resolution of genetic associations, and providing actionable insights for follow up work.

Genetics

Chromatin Landscapes and Genetic Risk For Juvenile Idiopathic Arthritis

Juvenile idiopathic arthritis (JIA) is considered to be an autoimmune disease mediated by interactions between genes and the environment. To gain a better understanding of the cellular basis for genetic risk, we studied known JIA genetic risk loci, the majority of which are located in non-coding regions, in human neutrophils and CD4 primary T cells to identify genes and functional elements located within those risk loci. We analyzed RNA-Seq data, H3K27ac and H3K4me1 chromatin immunoprecipitation-sequencing (ChIP-Seq) data, and previously published chromatin interaction analysis by paired-end tag sequencing (ChIA-PET) data to characterize the chromatin landscapes within the know JIA-associated risk loci. In both neutrophils and primary CD4+ T cells, the majority of the JIA-associated LD blocks contained H3K27ac and/or H3K4me1 marks. These LD blocks were also binding sites for a small group of transcription factors, particularly in neutrophils. Furthermore, these regions showed abundant intronic and intergenic transcription in neutrophils. In neutrophils, none of the genes that were differentially expressed between untreated JIA patients and healthy children was located within the JIA risk LD blocks. In CD4+ T cells, multiple genes, including HLA-DQA1, HLA-DQB2, TRAF1, and IRF1 were associated with the long-distance interacting regions within the LD regions as determined from ChIA-PET data. These findings suggest that aberrant transcriptional control is the underlying pathogenic mechanism in JIA. Furthermore, these findings demonstrate the challenges of identifying the actual causal variants within complex genomic/chromatin landscapes.

Genetics

Educational attainment and personality are genetically intertwined

Heritable variance in psychological traits may reflect genetic and biological processes that are not necessarily specific to these particular traits but pertain to a broader range of phenotypes. We tested the possibility that Five-Factor Model personality domains and their 30 facets, as rated by people themselves and their knowledgeable informants, reflect polygenic influences that have been previously associated with educational attainment. In a sample of over 3,000 adult Estonians, polygenic scores for educational attainment (EPS; interpretable as estimates of molecular genetic propensity for education) were correlated with various personality traits, particularly from the Neuroticism and Openness domains. The correlations of personality traits with phenotypic educational attainment closely mirrored their correlations with EPS. Moreover, EPS predicted an aggregate personality trait tailored to capture maximum amount of variance in educational attainment almost as strongly as it predicted the attainment itself. We discuss possible interpretations and implications of these findings.

Genetics

biMM: Efficient estimation of genetic variances andcovariances for cohorts with high-dimensional phenotype measurements

Genetic research utilizes a decomposition of trait variances and covariances into genetic and environmental parts. Our software package biMM is a computationally efficient implementation of a bivariate linear mixed model for settings where hundreds of traits have been measured on partially overlapping sets of individuals.\n\nAvailabilityImplementation in R freely available at www.iki.fi/mpirinen.

genetics

Genetic variation and gene expression across multiple tissues and developmental stages in a non-human primate

By analyzing multi-tissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalogue of expression quantitative trait loci (eQTLs) in a non-human primate model. This catalogue contains more genome-wide significant eQTLs, per sample, than comparable human resources, and reveals sex and age-related expression patterns. Findings include a master regulatory locus that likely plays a role in immune function, and a locus regulating hippocampal long non-coding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.

genetics

Transcriptome analysis of genetically matched human induced pluripotent stem cells disomic or trisomic for chromosome 21.

Trisomy of chromosome 21, the genetic cause of Down syndrome, has the potential to alter expression of genes on chromosome 21, as well as other locations throughout the genome. These transcriptome changes are likely to underlie the Down syndrome clinical phenotypes. We have employed RNA-seq to undertake an in-depth analysis of transcriptome changes resulting from trisomy of chromosome 21, using induced pluripotent stem cells (iPSCs) derived from a single individual with Down syndrome. These cells were originally derived by Li et al, who genetically targeted chromosome 21 in trisomic iPSCs, allowing selection of disomic sibling iPSC clones. Analyses were conducted on trisomic/disomic cell pairs maintained as iPSCs or differentiated into cortical neuronal cultures. In addition to characterization of gene expression levels, we have also investigated patterns of RNA adenosine-to-inosine editing, alternative splicing, and repetitive element expression, aspects of the transcriptome that have not been significantly characterized in the context of Down syndrome. We identified significant changes in transcript accumulation associated with chromosome 21 trisomy, as well as changes in alternative splicing and repetitive element transcripts. Unexpectedly, the trisomic iPSCs we characterized expressed higher levels of neuronal transcripts than control disomic iPSCs, and readily differentiated into cortical neurons, in contrast to another reported study. Comparison of our transcriptome data with similar studies of trisomic iPSCs suggests that trisomy of chromosome 21 may not intrinsically limit neuronal differentiation, but instead may interfere with the maintenance of pluripotency.

genetics

Genetic basis of melanin pigmentation in butterfly wings

Despite the variety, prominence, and adaptive significance of butterfly wing patterns surprisingly little known about the genetic basis of wing color diversity. Even though there is intense interest in wing pattern evolution and development, the technical challenge of genetically manipulating butterflies has slowed efforts to functionally characterize color pattern development genes. To identify candidate wing pigmentation genes we used RNA-seq to characterize transcription across multiple stages of butterfly wing development, and between different color pattern elements, in the painted lady butterfly Vanessa cardui. This allowed us to pinpoint genes specifically associated with red and black pigment patterns. To test the functions of a subset of genes associated with presumptive melanin pigmentation we used CRISPR/Cas9 genome editing in four different butterfly genera. pale, Ddc, and yellow knockouts displayed reduction of melanin pigmentation, consistent with previous findings in other insects. Interestingly, however, yellow-d, ebony, and black knockouts revealed that these genes have localized effects on tuning the color of red, brown, and ochre pattern elements. These results point to previously undescribed mechanisms for modulating the color of specific wing pattern elements in butterflies, and provide an expanded portrait of the insect melanin pathway.

genetics

The Genetic History of Northern Europe

Recent ancient DNA studies have revealed that the genetic history of modern Europeans was shaped by a series of migration and admixture events between deeply diverged groups. While these events are well described in Central and Southern Europe, genetic evidence from Northern Europe surrounding the Baltic Sea is still sparse. Here we report genome-wide DNA data from 24 ancient North Europeans ranging from [~]7,500 to 200 calBCE spanning the transition from a hunter-gatherer to an agricultural lifestyle, as well as the adoption of bronze metallurgy. We show that Scandinavia was settled after the retreat of the glacial ice sheets from a southern and a northern route, and that the first Scandinavian Neolithic farmers derive their ancestry from Anatolia 1000 years earlier than previously demonstrated. The range of Western European Mesolithic hunter-gatherers extended to the east of the Baltic Sea, where these populations persisted without gene-flow from Central European farmers until around 2,900 calBCE when the arrival of steppe pastoralists introduced a major shift in economy and established wide-reaching networks of contact within the Corded Ware Complex.

genetics

Genotype-phenotype association mining in bipolar disorder: market research meets complex genetics

Disentangling the etiology of common, complex diseases is a major challenge in genetic research. For bipolar disorder (BD), several genome-wide association studies (GWAS) have been performed. Similar to other complex disorders, major breakthroughs in explaining the high heritability of BD through GWAS have remained elusive. To overcome this dilemma, genetic research into BD, has embraced a variety of strategies such as the formation of large consortia to increase sample size and sequencing approaches. Here we advocate a complementary approach making use of already existing GWAS data: applying a data mining procedure to identify yet undetected genotype-phenotype relationships. We adapted association rule mining, a data mining technique traditionally used in retail market research, to identify frequent and characteristic genotype patterns showing strong associations to phenotype clusters. We applied this strategy to three independent GWAS datasets from 2,835 phenotypically characterized patients with BD. In a discovery step, 20,882 candidate association rules were extracted. Two of these - one associated with eating disorder and the other with anxiety - remained significant in an independent dataset after robust correction for multiple testing, showing considerable effect sizes (odds ratio ~ 3.4 and 3.0, respectively). Our approach may help detect novel specific genotype-phenotype relationships in BD typically not explored by analyses like GWAS. While we adapted the data mining tool within the context of BD gene discovery, it may facilitate identifying highly specific genotype-phenotype relationships in subsets of genome-wide data sets of other complex phenotype with similar epidemiological properties and challenges to gene discovery efforts.

genomics

Genetic control of age-related gene expression and complex traits in the human brain

Age is the primary risk factor for many of the most common human diseases--particularly neurodegenerative diseases--yet we currently have a very limited understanding of how each individuals genome affects the aging process. Here we introduce a method to map genetic variants associated with age-related gene expression patterns, which we call temporal expression quantitative trait loci (teQTL). We found that these loci are markedly enriched in the human brain and are associated with neurodegenerative diseases such as Alzheimers disease and Creutzfeldt-Jakob disease. Examining potential molecular mechanisms, we found that age-related changes in DNA methylation can explain some cis-acting teQTLs, and that trans-acting teQTLs can be mediated by microRNAs. Our results suggest that genetic variants modifying age-related patterns of gene expression, acting through both cis- and trans-acting molecular mechanisms, could play a role in the pathogenesis of diverse neurological diseases.

genetics

Natural Genetic Variation Can Independently Tune The Induced Fraction And Induction Level Of A Bimodal Signaling Response

Bimodal gene expression by genetically identical cells is a pervasive feature of signaling networks. In the galactose-utilization (GAL) pathway of Saccharomyces cerevisiae, induction can be unimodal or bimodal depending on natural genetic variation and pre-induction conditions. Here, we find that this variation of modality is regulated by an interplay between two features of the pathway response, the fraction of cells that are in the induced subpopulation and their expression level. Combined, the variations in these features are sufficient to explain the observed effects of natural variation and pre-induction conditions on the modality of induction in both mechanistic and phenomenological models. Both natural variation and pre-induction conditions act by modulating the expression and function of the galactose sensor GAL3. The ability to alter modality may allow organisms to adapt their level of "bet hedging" to the conditions they experience, and thus help optimize fitness in complex, fluctuating natural environments.

genetics

Genetic Instrumental Variable (GIV) Regression: Explaining Socioeconomic And Health Outcomes In Non-Experimental Data

Identifying causal effects in non-experimental data is an enduring challenge. One proposed solution that recently gained popularity is the idea to use genes as instrumental variables (i.e. Mendelian Randomization - MR). However, this approach is problematic because many variables of interest are genetically correlated, which implies the possibility that many genes could affect both the exposure and the outcome directly or via unobserved confounding factors. Thus, pleiotropic effects of genes are themselves a source of bias in non-experimental data that would also undermine the ability of MR to correct for endogeneity bias from non-genetic sources. Here, we propose an alternative approach, GIV regression, that provides estimates for the effect of an exposure on an outcome in the presence of pleiotropy. As a valuable byproduct, GIV regression also provides accurate estimates of the chip heritability of the outcome variable. GIV regression uses polygenic scores (PGS) for the outcome of interest which can be constructed from genome-wide association study (GWAS) results. By splitting the GWAS sample for the outcome into non-overlapping subsamples, we obtain multiple indicators of the outcome PGS that can be used as instruments for each other, and, in combination with other methods such as sibling fixed effects, can address endogeneity bias from both pleiotropy and the environment. In two empirical applications, we demonstrate that our approach produces reasonable estimates of the chip heritability of educational attainment (EA) and show that standard regression and MR provide upwardly biased estimates of the effect of body height on EA.

genetics

The Stratification Of Major Depressive Disorder Into Genetic Subgroups

Depression is a common and clinically heterogeneous mental health disorder that is frequently comorbid with other diseases and conditions. Stratification of depression may align sub-diagnoses more closely with their underling aetiology and provide more tractable targets for research and effective treatment. In the current study, we investigated whether genetic data could be used to identify subgroups within people with depression using the UK Biobank. Examination of cross-locus correlations was used to test for evidence of subgroups by examining whether there was clustering of independent genetic variants associated with eleven other complex traits and disorders in people with depression. We found evidence of a subgroup within depression using age of natural menopause variants (P = 1.69 x 10-3) and this effect remained significant in females (P = 1.18 x 10-3), but not males (P = 0.186). However, no evidence for this subgroup (P > 0.05) was found in Generation Scotland, iPSYCH, a UK Biobank replication cohort or the GERA cohort. In the UK Biobank, having depression was also associated with a later age of menopause (beta = 0.34, standard error = 0.06, P = 9.92 x 10-8). A potential age of natural menopause subgroup within depression and the association between depression and a later age of menopause suggests that they partially share a developmental pathway.

genetics

Genome-Wide Association Study Reveals Genetic Link Between Diarrhea-Associated Entamoeba histolytica Infection And Inflammatory Bowel Disease

Diarrhea is the second leading cause of death for children globally, causing 760,000 deaths each year in children under the age of 5. Amoebic dysentery contributes significantly to this burden, especially in developing countries. We hypothesize that genetic variation contributes to susceptibility to diarrhea-associated Entamoeba histolytica infection in Bangladeshi infants; thus, we conducted a genome-wide association study (GWAS) in two independent birth cohorts of diarrhea-associated E. histolytica infection. Cases were defined as children with at least one diarrheal episode positive for E. histolytica through either PCR or ELISA within the first year of life. Controls were children without any episodes positive for E. histolytica in the same time frame. Meta-analyses under a fixed-effects inverse variance weighting model identified variants in two neighboring genes on chromosome 10: CUL2 (cullin 2) and CREM (cAMP responsive element modulator) associated with E. histolytica infection, with SNP rs58000832 achieving genome-wide significance (Pmeta=4.2x10-10). Each additional risk allele (an intergenic insertion between CREM and CCNY) of rs58000832 conferred 2.5 increased odds of a diarrhea-associated E. histolytica infection. The most associated SNP within a gene was in an intron of CREM (rs58468685, Pmeta=2.3x10-9), which with CUL2, has been implicated as a susceptibility locus for Inflammatory Bowel Disease (IBD) and Crohns Disease. Gene expression resources suggest these loci are related to the higher expression of CREM, but not CUL2. Increased CREM expression is also observed in early E. histolytica infection. Further, CREM-/- mice were more susceptible to E. histolytica amebic colitis. These genetic associations reinforce the pathological similarities observed in gut inflammation between E. histolytica infection and IBD.

genetics

The iPSYCH2012 case-cohort sample: New directions for unravelling genetic and environmental architectures of severe mental disorders

The iPSYCH consortium has established a large Danish population-based Case-Cohort sample (iPSYCH2012) aimed at unravelling the genetic and environmental architecture of severe mental disorders. The iPSYCH2012 sample is nested within the entire Danish population born 1981-2005 including 1,472,762 persons. This paper introduces the iPSYCH2012 sample and outlines key future research directions. Cases were identified as persons with schizophrenia (N=3,540), autism (N=16,146), ADHD (N=18,726), and affective disorder (N=26,380), of which 1928 had bipolar affective disorder. Controls were randomly sampled individuals (N=30,000). Within the sample of 86,189 individuals, a total of 57,377 individuals had at least one major mental disorder. DNA was extracted from the neonatal dried blood spot samples obtained from the Danish Neonatal Screening Biobank and genotyped using the Illumina PsychChip. Genotyping was successful for 90% of the sample. The assessments of exome sequencing, methylation profiling, metabolome profiling, vitamin-D, inflammatory and neurotrophic factors are in progress. For each individual, the iPSYCH2012 sample also includes longitudinal information on health, prescribed medicine, social and socioeconomic information and analogous information among relatives. To the best of our knowledge, the iPSYCH2012 sample is the largest and most comprehensive data source for the combined study of genetic and environmental aetiologies of severe mental disorders.

genetics

Comprehensive genome and transcriptome analysis reveals genetic basis for gene fusions in cancer

Gene fusions are an important class of cancer-driving events with therapeutic and diagnostic values, yet their underlying genetic mechanisms have not been systematically characterized. Here by combining RNA and whole genome DNA sequencing data from 1188 donors across 27 cancer types we obtained a list of 3297 high-confidence tumour-specific gene fusions, 82% of which had structural variant (SV) support and 2372 of which were novel. Such a large collection of RNA and DNA alterations provides the first opportunity to systematically classify the gene fusions at a mechanistic level. While many could be explained by single SVs, numerous fusions involved series of structural rearrangements and thus are composite fusions. We discovered 75 fusions of a novel class of inter-chromosomal composite fusions, termed bridged fusions, in which a third genomic location bridged two different genes. In addition, we identified 522 fusions involving non-coding genes and 157 ORF-retaining fusions, in which the complete open reading frame of one gene was fused to the UTR region of another. Although only a small proportion (5%) of the discovered fusions were recurrent, we found a set of highly recurrent fusion partner genes, which exhibited strong 5 or 3 bias and were significantly enriched for cancer genes. Our findings broaden the view of the gene fusion landscape and reveal the general properties of genetic alterations underlying gene fusions for the first time.

genetics

Genome-wide analysis reveals distinct genetic mechanisms of diet-dependent lifespan and healthspan in D. melanogaster

Dietary restriction (DR) robustly extends lifespan and delays age-related diseases across species. An underlying assumption in aging research has been that DR mimetics extend both lifespan and healthspan jointly, though this has not been rigorously tested in different genetic backgrounds. Furthermore, nutrient response genes important for lifespan or healthspan extension remain underexplored, especially in natural populations. To address these gaps, we utilized over 150 DGRP strains to measure nutrient-dependent changes in lifespan and age-related climbing ability to measure healthspan. DR extended lifespan and delayed decline in climbing ability on average, but there was no evidence of correlation between these traits across individual strains. Through GWAS, we then identified and validated jughead and Ferredoxin as determinants of diet-dependent lifespan, and Daedalus for diet-dependent physical activity. Modulating these genes produced independent effects on lifespan and climbing ability, further suggesting that these age-related traits are likely to be regulated through distinct genetic mechanisms.

genetics