bioRxiv ScienceSearch

Biology subjects

Jo Knight

Publications and source records attributed to Jo Knight.

8 recordsLinked to original sources

Semi-Automated Identification of Ontological Labels in the Biomedical Literature with goldi

Recent growth in both the scale and the scope of large publicly available ontologies has spurred the development of computational methodologies which can leverage structured information to answer important questions. However, ontological labels, or \"terms\" have thus far proved difficult to use in practice; text mining, one crucial aspect of electronically understanding and parsing the biomedical literature, has historically had difficulty identifying \"terms\" in literature. In this article, we present goldi, an open source R package whose goal it is to identify terms of variable length in free form text. It is available at https://github.com/Chris1221/goldi or through CRAN. The algorithm works through identifying words or synonyms of words present in individual terms and comparing the number of present words to an acceptance function for decision making. In this article we present the theoretical rationale behind the algorithm, as well as practical advice for its usage applied to Gene Ontology term identification and quantification. We additionally detail the options available and describe their respective computational efficiencies.

Bioinformatics

Polygenic analysis of schizophrenia and 19 immune diseases reveals modest pleiotropy

Epidemiological studies indicate that many immune diseases occur at different rates among people with schizophrenia compared to the general population. Here, we evaluated whether this phenotypic correlation between immune diseases and schizophrenia might be explained by shared genetic risk factors (genetic correlation). We used data from a large genome-wide association study (GWAS) of schizophrenia (N=35,476 cases and 46,839 controls) to compare the genetic architecture of schizophrenia to 19 immune diseases. First, we evaluated the association with schizophrenia of 581 variants previously reported to be associated with immune diseases at genome-wide significance. We identified three variants with pleiotropic effects, located in regions associated with both schizophrenia and immune disease. Our analyses provided the strongest evidence of pleiotropy at rs1734907 ([~]85kb upstream of EPHB4), a variant which was associated with increased risk of both Crohns disease (OR = 1.16, P = 1.67x10-13) and schizophrenia (OR = 1.07, P = 7.55x10-6). Next, we investigated genome-wide sharing of common variants between schizophrenia and immune diseases using polygenic risk scores (PRS) and cross-trait LD Score regression (LDSC). PRS revealed significant genetic overlap with schizophrenia for narcolepsy (p=4.1x10-4), primary biliary cirrhosis (p=1.4x10-8), psoriasis (p=3.6x10-5), systemic lupus erythematosus (p=2.2x10-8), and ulcerative colitis (p=4.3x10-4). Genetic correlations between these immune diseases and schizophrenia, estimated using LDSC, ranged from 0.10 to 0.18 and were consistent with the expected phenotypic correlation based on epidemiological data. We also observed suggestive evidence of sex-dependent genetic correlation between schizophrenia and multiple sclerosis (interaction p=0.02), with genetic risk scores for multiple sclerosis associated with greater risk of schizophrenia among males but not females. Our findings suggest that shared genetic risk factors contribute to the epidemiological co-occurrence of schizophrenia and certain immune diseases, and suggest that in some cases this genetic correlation is sex-dependent. Author SummaryImmune diseases occur at different rates among patients with schizophrenia compared to the general population. While the reasons for this phenotypic correlation are unclear, shared genetic risk (genetic correlation) has been proposed as a contributing factor. Prior studies have estimated the genetic correlation between schizophrenia and a handful of immune diseases, with conflicting results. Here, we performed a comprehensive cross-disorder investigation of schizophrenia and 19 immune diseases. We identified three individual genetic variants associated with both schizophrenia and immune diseases, including a variant near EPHB4 - a gene whose protein product guides the migration of lymphocytes towards infected cells in the immune system and the migration of neuronal axons in the brain. We demonstrated significant genome-wide genetic correlation between schizophrenia and narcolepsy, primary biliary cirrhosis, psoriasis, systemic lupus erythematosus, and ulcerative colitis. Finally, we identified a potential sex-dependent pleiotropic effect between schizophrenia and multiple sclerosis. Our findings point to shared genetic risk for schizophrenia and at least a subset of immune diseases, which likely contributes to their epidemiological co-occurrence. These results raise the possibility that the same genetic variants may exert their effects on neurons or immune cells to influence the development of psychiatric and immune disorders, respectively.

Genetics

Genetic variability in both the adaptive and innate immune systems contribute to Alzheimer’s and Parkinson’s disease risk

Neurodegenerative disorders are devastating diseases with a worldwide health-care burden. Studies have demonstrated enrichment of disease-associated genetic variants with functional genomic annotations. Determining associated cell-types is important to understand pathogenicity.\n\nWe obtained GWAS summary statistics from Parkinsons disease (PD), Alzheimers disease (AD), amyotrophic lateral sclerosis (ALS), multiple sclerosis (MS), and frontotemporal dementia (FTD). We applied stratified LD score regression to determine if functional categories are enriched for heritability.\n\nThere was little enrichment of brain annotations, but annotations from both the innate and adaptive immune systems were enriched for MS (as expected), AD, and PD, in decreasing order of statistical significance.

Genomics

Genome-wide association studies suggest limited immune gene enrichment in schizophrenia compared to five autoimmune diseases

There has been intense debate over the immunological basis of schizophrenia, and the potential utility of adjunct immunotherapies. The major histocompatibility complex is consistently the most powerful region of association in genome-wide association studies (GWASs) of schizophrenia, and has been interpreted as strong genetic evidence supporting the immune hypothesis. However, global pathway analyses provide inconsistent evidence of immune involvement in schizophrenia, and it remains unclear whether genetic data support an immune etiology per se. Here we empirically test the hypothesis that variation in immune genes contributes to schizophrenia. We show that there is no enrichment of immune loci outside of the MHC region in the largest genetic study of schizophrenia conducted to date, in contrast to five diseases of known immune origin. Among 108 regions of the genome previously associated with schizophrenia, we identify six immune candidates (DPP4, HSPD1, EGR1, CLU, ESAM, NFATC3) encoding proteins with alternative, nonimmune roles in the brain. While our findings do not refute evidence that has accumulated in support of the immune hypothesis, they suggest that genetically mediated alterations in immune function may not play a major role in schizophrenia susceptibility. Instead, there may be a role for pleiotropic effects of a small number of immune genes that also regulate brain development and plasticity. Whether immune alterations drive schizophrenia progression is an important question to be addressed by future research, especially in light of the growing interest in applying immunotherapies in schizophrenia.

Genetics

Association mapping of inflammatory bowel disease loci to single variant resolution

Inflammatory bowel disease (IBD) is a chronic gastrointestinal inflammatory disorder that affects millions worldwide. Genome-wide association studies (GWAS) have identified 200 IBD-associated loci, but few have been conclusively resolved to specific functional variants. Here we report fine-mapping of 94 IBD loci using high-density genotyping in 67,852 individuals. Of the 139 independent associations identified in these regions, 18 were pinpointed to a single causal variant with >95% certainty, and an additional 27 associations to a single variant with >50% certainty. These 45 variants are significantly enriched for protein-coding changes (n=13), direct disruption of transcription factor binding sites (n=3) and tissue specific epigenetic marks (n=10), with the latter category showing enrichment in specific immune cells among associations stronger in CD and gut mucosa among associations stronger in UC. The results of this study suggest that high-resolution, fine-mapping in large samples can convert many GWAS discoveries into statistically convincing causal variants, providing a powerful substrate for experimental elucidation of disease mechanisms.

Genetics

Novel bioinformatics approach to investigate quantitative phenotype-genotype associations in neuroimaging studies

Imaging genetics is an emerging field in which the association between genes and neuroimaging-based quantitative phenotypes are used to explore the functional role of genes in neuroanatomy and neurophysiology in the context of healthy function and neuropsychiatric disorders. The main obstacle for researchers in the field is the high dimensionality of the data in both the imaging phenotypes and the genetic variants commonly typed. In this article, we develop a novel method that utilizes Gene Ontology, an online database, to select and prioritize certain genes, employing a stratified false discovery rate (sFDR) approach to investigate their associations with imaging phenotypes. sFDR has the potential to increase power in genome wide association studies (GWAS), and is quickly gaining traction as a method for multiple testing correction. Our novel approach addresses both the pressing need in genetic research to move beyond candidate gene studies, while not being overburdened with a loss of power due to multiple testing. As an example of our methodology, we perform a GWAS of hippocampal volume using the Alzheimers Disease Neuroimaging Initiative sample.

Genetics

Circumstantial Evidence? Comparison of Statistical Learning Methods using Functional Annotations for Prioritizing Risk Variants

Although technology has triumphed in facilitating routine genome re-sequencing, new challenges have been created for the data analyst. Genome scale surveys of human disease variation generate volumes of data that far exceed capabilities for laboratory characterization, and importantly also create a substantial burden of type I error. By incorporating a variety of functional annotations as predictors, such as regulatory and protein coding elements, statistical learning has been widely investigated as a mechanism for the prioritization of genetic variants that are more likely to be associated with complex disease. These methods offer a hope of identification of sufficiently large numbers of truly associated variants, to make cost-effective the large-scale functional characterization necessary to progress genome scale experiments. We compared the results from three published prioritization procedures which use different statistical learning algorithms and different predictors with regard to the quantity, type and coding of the functional annotations. In this paper we also explore different combinations of algorithm and annotation set. We train the models in 60% of the data and reserve the remainder for testing the accuracy. As an application, we tested which methodology performed the best for prioritizing sub-genome-wide-significant variants (5x10-8<p<1x10-6) using data from the first and second rounds of a large schizophrenia meta-analysis by the Psychiatric Genomics Consortium. Results suggest that all methods have considerable (and similar) predictive accuracies (AUCs 0.64-0.71). However, predictive accuracy results obtained from the test set do not always reflect results obtained from the application to the schizophrenia meta-analysis. In conclusion, a variety of algorithms and annotations seem to have a similar potential to effectively enrich true risk variants in genome scale datasets, however none offer more than incremental improvement in prediction. We discuss how methods might be evolved towards the step change in the risk variant prediction required to address the impending bottleneck of the new generation of genome re-sequencing studies.

Bioinformatics

A Bayesian Method to Incorporate Hundreds of Functional Characteristics with Association Evidence to Improve Variant Prioritization

The increasing quantity and quality of functional genomic information motivate the assessment and integration of these data with association data, including data originating from genome-wide association studies (GWAS). We used previously described GWAS signals (\"hits\") to train a regularized logistic model in order to predict SNP causality on the basis of a large multivariate functional dataset. We show how this model can be used to derive Bayes factors for integrating functional and association data into a combined Bayesian analysis. Functional characteristics were obtained from the Encyclopedia of DNA Elements (ENCODE), from published expression quantitative trait loci (eQTL), and from other sources of genome-wide characteristics. We trained the model using all GWAS signals combined, and also using phenotype specific signals for autoimmune, brain-related, cancer, and cardiovascular disorders. The non-phenotype specific and the autoimmune GWAS signals gave the most reliable results. We found SNPs with higher probabilities of causality from functional characteristics showed an enrichment of more significant p-values compared to all GWAS SNPs in three large GWAS studies of complex traits. We investigated the ability of our Bayesian method to improve the identification of true causal signals in a psoriasis GWAS dataset and found that combining functional data with association data improves the ability to prioritise novel hits. We used the predictions from the penalized logistic regression model to calculate Bayes factors relating to functional characteristics and supply these online alongside resources to integrate these data with association data.\n\nAuthor SummaryLarge-scale genetic studies have had success identifying genes that play a role in complex traits. Advanced statistical procedures suggest that there are still genetic variants to be discovered, but these variants are difficult to detect. Incorporating biological information that affect the amount of protein or other product produced can be used to prioritise the genetic variants in order to identify which are likely to be causal. The method proposed here uses such biological characteristics to predict which genetic variants are most likely to be causal for complex traits.

Bioinformatics