bioRxiv ScienceSearch

Biology subjects

Flint, J.

Publications and source records attributed to Flint, J..

10 recordsLinked to original sources

A comprehensive analysis of the usability and archival stability of omics computational tools and resources

Developing new software tools for analysis of large-scale biological data is a key component of advancing modern biomedical research. Scientific reproduction of published findings requires running computational tools on data generated by such studies, yet little attention is presently allocated to the installability and archival stability of computational software tools. Scientific journals require data and code sharing, but none currently require authors to guarantee the continuing functionality of newly published tools. We have estimated the archival stability of computational biology software tools by performing an empirical analysis of the internet presence for 36,702 omics software resources published from 2005 to 2017. We found that almost 28% of all resources are currently not accessible through URLs published in the paper they first appeared in. Among the 98 software tools selected for our installability test, 51% were deemed \"easy to install,\" and 28% of the tools failed to be installed at all due to problems in the implementation. Moreover, for papers introducing new software, we found that the number of citations significantly increased when authors provided an easy installation process. We propose for incorporation into journal policy several practical solutions for increasing the widespread installability and archival stability of published bioinformatics software.

bioinformatics

Reverse GWAS: Using Genetics to Identify and Model Phenotypic Subtypes

Recent and classical work has revealed biologically and medically significant subtypes in complex diseases and traits. However, relevant subtypes are often unknown, unmeasured, or actively debated, making automatic statistical approaches to subtype definition particularly valuable. We propose reverse GWAS (RGWAS) to identify and validate subtypes using genetics and multiple traits: while GWAS seeks the genetic basis of a given trait, RGWAS seeks to define trait subtypes with distinct genetic bases. Unlike existing approaches relying on off-the-shelf clustering methods, RGWAS uses a bespoke decomposition, MFMR, to model covariates, binary traits, and population structure. We use extensive simulations to show these features can be crucial for power and calibration. We validate RGWAS in practice by recovering known stress subtypes in major depressive disorder. We then show the utility of RGWAS by identifying three novel subtypes of metabolic traits. We biologically validate these metabolic subtypes with SNP-level tests and a novel polygenic test: the former recover known metabolic GxE SNPs; the latter suggests genetic heterogeneity may explain substantial missing heritability. Crucially, statins, which are widely prescribed and theorized to increase diabetes risk, have opposing effects on blood glucose across metabolic subtypes, suggesting potential have potential translational value.\n\nAuthor summaryComplex diseases depend on interactions between many known and unknown genetic and environmental factors. However, most studies aggregate these strata and test for associations on average across samples, though biological factors and medical interventions can have dramatically different effects on different people. Further, more-sophisticated models are often infeasible because relevant sources of heterogeneity are not generally known a priori. We introduce Reverse GWAS to simultaneously split samples into homogeneoues subtypes and to learn differences in genetic or treatment effects between subtypes. Unlike existing approaches to computational subtype identification using high-dimensional trait data, RGWAS accounts for covariates, binary disease traits and, especially, population structure; these features are each invaluable in extensive simulations. We validate RGWAS by recovering known genetic subtypes of major depression. We demonstrate RGWAS is practically useful in a metabolic study, finding three novel subtypes with both SNP- and polygenic-level heterogeneity. Importantly, RGWAS can uncover differential treatment response: for example, we show that statin, a common drug and potential type 2 diabetes risk factor, may have opposing subtype-specific effects on blood glucose.

genetics

Minimal phenotyping yields GWAS hits of low specificity for major depression

Minimal phenotyping refers to the reliance on the use of a small number of self-report items for disease case identification. This strategy has been applied to genome-wide association studies (GWAS) of major depressive disorder (MDD). Here we report that the genotype derived heritability (h2SNP) of depression defined by minimal phenotyping (14%, SE = 0.8%) is lower than strictly defined MDD (26%, SE = 2.2%). This cannot be explained by differences in prevalence between definitions or including cases of lower liability to MDD in minimal phenotyping definitions of depression, but can be explained by misdiagnosis of those without depression or with related conditions as cases of depression. Depression defined by minimal phenotyping is as genetically correlated with strictly defined MDD (rG = 0.81, SE = 0.03) as it is with the personality trait neuroticism (rG = 0.84, SE = 0.05), a trait not defined by the cardinal symptoms of depression. While they both show similar shared genetic liability with neuroticism, a greater proportion of the genome contributes to the minimal phenotyping definitions of depression (80.2%, SE = 0.6%) than to strictly defined MDD (65.8%, SE = 0.6%). We find that GWAS loci identified in minimal phenotyping definitions of depression are not specific to MDD: they also predispose to other psychiatric conditions. Finally, while highly predictive polygenic risk scores can be generated from minimal phenotyping definitions of MDD, the predictive power can be explained entirely by the sample size used to generate the polygenic risk score, rather than specificity for MDD. Our results reveal that genetic analysis of minimal phenotyping definitions of depression identifies non-specific genetic factors shared between MDD and other psychiatric conditions. Reliance on results from minimal phenotyping for MDD may thus bias views of the genetic architecture of MDD and may impede our ability to identify pathways specific to MDD.

genetics

GxEMM: Extending linear mixed models to general gene-environment interactions

Gene-environment interaction (GxE) is a well-known source of non-additive inheritance. GxE can be important in applications ranging from basic functional genomics to precision medical treatment. Further, GxE effects elude inherently-linear LMMs and may explain missing heritability. We propose a simple, unifying mixed model for polygenic interactions (GxEMM) to capture the aggregate effect of small GxE effects spread across the genome. GxEMM extends existing LMMs for GxE in two important ways. First, it extends to arbitrary environmental variables, not just categorical groups. Second, GxEMM can estimate and test for environment-specific heritability. In simulations where the assumptions of existing methods do not hold, we show that GxEMM improves estimates of ordinary and GxE heritability and increases power to test for polygenic GxE. We then use GxEMM to prove that the heritability of major depression (MD) is reduced by stress, which we previously conjectured but could not prove with prior methods, and that a tail of polygenic GxE effects remains unexplained by MD GWAS.

genetics

JEPEGMIX2-P: Novel pathway transcriptomic method greatly increases detection of molecular pathways in cosmopolitan cohorts

Genetic signal detection in genome-wide association studies (GWAS) is enhanced by pooling small signals from multiple Single Nucleotide Polymorphism (SNP), e.g. across genes and pathways. Because genes are believed to influence traits via gene expression, it is of interest to combine information from expression Quantitative Trait Loci (eQTLs) in a gene or genes in the same pathway. Such methods, widely referred as transcriptomic wide association analysis (TWAS), already exist for gene analysis. Due to the possibility of eliminating most of the confounding effect of linkage disequilibrium (LD) from TWAS gene statistics, pathway TWAS methods would be very useful in uncovering the true molecular bases of psychiatric disorders. However, such methods are not yet available for arbitrarily large pathways/gene sets. This is possibly due to it quadratic (in the number of SNPs) computational burden for computing LD across large regions. To overcome this obstacle, we propose JEPEGMIX2-P, a novel TWAS pathway method that i) has a linear computational burden, ii) uses a large and diverse reference panel (33K subjects), iii) is competitive (adjusts for background enrichment in gene TWAS statistics) and iv) is applicable as-is to ethnically mixed cohorts. To underline its potential for increasing the power to uncover genetic signals over the state-of-the-art and commonly used non-transcriptomics methods, e.g. MAGMA, we applied JEPEGMIX2-P to summary statistics of most large meta-analyses from Psychiatric Genetics Consortium (PGC). While our work is just the very first step toward clinical translation of psychiatric disorders, PGC anorexia results suggest a possible avenue for treatment.

bioinformatics

Coping-Style Behaviour Identified by a Survey of Parent-of-Origin Effects in the Rat

We develop theory, based on earlier work, to partition heritability into a component due to a combination of parent of origin, maternal, paternal and shared environment, and another component that estimates classical additive genetic variance. We then investigate the effects on heritability of the parental origin of alleles in outbred heterogeneous stock rats across 199 complex traits. Parent-of-origin-like heritability was on average 2.7-fold larger than classical additive heritability. Among the phenotypes with the most enhanced parent-of-origin heritability were 10 coping style behaviors, with average 3.2-fold heritability enrichment. To confirm these findings on coping behaviour, and to eliminate the possibility that the parent of origin effects are due to confounding with shared environment, we performed a reciprocal F1 cross between the behaviourally divergent RHA and RLA rat strains. We observed parent-of-origin effects on F1 rat anxiety/coping-related behavior in the Elevated Zero Maze test. Our results are the first to assess genetic parent-of-origin effects in rats, and confirm earlier findings in mice that such effects influence mammalian coping and impulsive behavior.

genetics

Multiple laboratory mouse reference genomes define strain specific haplotypes and novel functional loci

The most commonly employed mammalian model organism is the laboratory mouse. A wide variety of genetically diverse inbred mouse strains, representing distinct physiological states, disease susceptibilities, and biological mechanisms have been developed over the last century. We report full length draft de novo genome assemblies for 16 of the most widely used inbred strains and reveal for the first time extensive strain-specific haplotype variation. We identify and characterise 2,567 regions on the current Genome Reference Consortium mouse reference genome exhibiting the greatest sequence diversity between strains. These regions are enriched for genes involved in defence and immunity, and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. Several immune related loci, some in previously identified QTLs for disease response have novel haplotypes not present in the reference that may explain the phenotype. We used these genomes to improve the mouse reference genome resulting in the completion of 10 new gene structures, and 62 new coding loci were added to the reference genome annotation. Notably this high quality collection of genomes revealed a previously unannotated gene (Efcab3-like) encoding 5,874 amino acids, one of the largest known in the rodent lineage. Interestingly, Efcab3-like-/- mice exhibit severe size anomalies in four regions of the brain suggesting a mechanism of Efcab3-like regulating brain development.

genomics

Estimating Heritability and Genetic Correlation in Case Control Studies Directly and with Summary Statistics

Methods that estimate heritability and genetic correlations from genome-wide association studies have proven to be powerful tools for investigating the genetic architecture of common diseases and exposing unexpected relationships between disorders. Many relevant studies employ a case-control design, yet most methods are primarily geared towards analyzing quantitative traits. Here we investigate the validity of three common methods for estimating genetic heritability and genetic correlation. We find that the Phenotype-Correlation-Genotype-Correlation (PCGC) approach is the only method that can estimate both quantities accurately in the presence of important non-genetic risk factors, such as age and sex. We extend PCGC to work with summary statistics that take the case-control sampling into account, and demonstrate that our new method, PCGC-s, accurately estimates both heritability and genetic correlations and can be applied to large data sets without requiring individual-level genotypic or phenotypic information. Finally, we use PCGC-S to estimate the genetic correlation between schizophrenia and bipolar disorder, and demonstrate that previous estimates are biased due to incorrect handling of sex as a strong risk factor. PCGC-s is available at https://github.com/omerwe/PCGCs.

genetics

Pathway-based polygenic risk implicates GO: 17144 drug metabolism in recurrent depressive disorder

The Psychiatric Genomics Consortium (PGC) has made major advances in the molecular etiology of MDD, confirming that MDD is highly polygenic, with any top risk loci conferring a very small proportion of variance in case-control status (1). Pathway enrichment results from PGC meta-analyses can also be used to help inform molecular drug targets. Prior to any knowledge of molecular biomarkers for MDD, drugs targeting molecular pathways have proved successful in treating MDD. However, it is possible that with information from PGC analyses, examining specific molecular pathway(s) implicated in MDD can further inform our study of molecular drug targets. Using a large case-control GWAS based on low-coverage whole genome sequencing (N = 10,640), we derived polygenic risk scores for MDD and for MDD specific to each of over 300 molecular pathways. We then used these data to identify sets of scores significantly predictive of case status, accounting for critical covariates. Over and above global polygenic risk for MDD, polygenic risk within the GO: 17144 drug metabolism pathway significantly predicted recurrent depression. In transcriptomic analyses, two pathway genes yielded suggestive signals at FDR q-values = .054: CYP2C19 (family of Cytochrome P450) and CBR1 (Carbonyl Reductase 1). Because the neuroleptic carbamazepine is a known inducer of CYP2C19, future research might examine whether drug metabolism PRS has any influence on clinical presentation and treatment response. Overall, results indicate that pathway-based risk might inform treatment of severe depression. We discuss limitations to the generalizability of these preliminary findings, and urge replication in future research.

genetics

A comprehensive map of genetic variation in the world’s largest ethnic group - Han Chinese

As are most non-European populations around the globe, the Han Chinese are relatively understudied in population and medical genetics studies. From low-coverage whole-genome sequencing of 11,670 Han Chinese women we present a catalog of 25,057,223 variants, including 548,401 novel variants that are seen at least 10 times in our dataset. Individuals from our study come from 19 out of 22 provinces across China, allowing us to study population structure, genetic ancestry, and local adaptation in Han Chinese. We identify previously unrecognized population structure along the East-West axis of China and report unique signals of admixture across geographical space, such as European influences among the Northwestern provinces of China. Finally, we identified a number of highly differentiated loci, indicative of local adaptation in the Han Chinese. In particular, we detected extreme differentiation among the Han Chinese at MTHFR, ADH7, and FADS loci, suggesting that these loci may not be specifically selected in Tibetan and Inuit populations as previously suggested. On the other hand, we find that Neandertal ancestry does not vary significantly across the provinces, consistent with admixture prior to the dispersal of modern Han Chinese. Furthermore, contrary to a previous report, Neandertal ancestry does not explain a significant amount of heritability in depression. Our findings provide the largest genetic data set so far made available for Han Chinese and provide insights into the history and population structure of the worlds largest ethnic group.

genetics