bioRxiv ScienceSearch

Biology subjects

Krumsiek, J.

Publications and source records attributed to Krumsiek, J..

4 recordsLinked to original sources

MoDentify: a tool for phenotype-driven module identification in multilevel metabolomics networks

SummaryMetabolomics is an established tool to gain insights into (patho)physiological outcomes. Associations of metabolism with such outcomes are expected to span functional modules, which are defined as sets of correlating metabolites that are coordinately regulated. Moreover, these associations occur at different scales, from entire pathways to only a few metabolites, which is an aspect that has not been addressed by previous methods. Here we present MoDentify, a freely available R package to identify regulated modules in metabolomics networks at different layers of resolution. Importantly, MoDentify shows higher statistical power than classical association analysis. Moreover, the package offers direct visualization of results as interactive networks in Cytoscape. We present an application example using a complex, multifluid metabolomics dataset. Owing to its generic character, the method is widely applicable to any dataset with a phenotype variable, a data matrix, and optional pathway annotations.\n\nAvailability and ImplementationMoDentify is freely available from GitHub: https://github.com/krumsiek/MoDentify\n\nThe package vignette contains a detailed tutorial of the analysis workflow.\n\nContactjan.krumsiek@helmholtz-muenchen.de

systems biology

Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies

BACKGROUNDUntargeted mass spectrometry (MS)-based metabolomics data often contain missing values that reduce statistical power and can introduce bias in epidemiological studies. However, a systematic assessment of the various sources of missing values and strategies to handle these data has received little attention. Missing data can occur systematically, e.g. from run day-dependent effects due to limits of detection (LOD); or it can be random as, for instance, a consequence of sample preparation.\n\nMETHODSWe investigated patterns of missing data in an MS-based metabolomics experiment of serum samples from the German KORA F4 cohort (n = 1750). We then evaluated 31 imputation methods in a simulation framework and biologically validated the results by applying all imputation approaches to real metabolomics data. We examined the ability of each method to reconstruct biochemical pathways from data-driven correlation networks, and the ability of the method to increase statistical power while preserving the strength of established genetically metabolic quantitative trait loci.\n\nRESULTSRun day-dependent LOD-based missing data accounts for most missing values in the metabolomics dataset. Although multiple imputation by chained equations (MICE) performed well in many scenarios, it is computationally and statistically challenging. K-nearest neighbors (KNN) imputation on observations with variable pre-selection showed robust performance across all evaluation schemes and is computationally more tractable.\n\nCONCLUSIONMissing data in untargeted MS-based metabolomics data occur for various reasons. Based on our results, we recommend that KNN-based imputation is performed on observations with variable pre-selection since it showed robust results in all evaluation schemes.\n\nKey messagesO_LIUntargeted MS-based metabolomics data show missing values due to both batch-specific LOD-based and non-LOD-based effects.\nC_LIO_LIStatistical evaluation of multiple imputation methods was conducted on both simulated and real datasets.\nC_LIO_LIBiological evaluation on real data assessed the ability of imputation methods to preserve statistical inference of biochemical pathways and correctly estimate effects of genetic variants on metabolite levels.\nC_LIO_LIKNN-based imputation on observations with variable pre-selection and K = 10 showed robust performance for all data scenarios across all evaluation schemes.\nC_LI

systems biology

Common patterns of gene regulation associated with Cesarean section and the development of islet autoimmunity -- indications of immune cell activation

BackgroundBirth by Cesarean section increases the risk of developing type 1 diabetes later in life; however, the underlying molecular mechanisms of this effect remain unclear. We aimed to elucidate common regulatory processes observed after Cesarean section and the development of islet autoimmunity, which precedes type 1 diabetes, by investigating the transcriptome of blood cells in the developing immune system.\n\nMethodsWe analyzed gene expression of peripheral blood mononuclear cells taken at several time points from children with increased familial and genetic risk for type 1 diabetes (n = 109). We investigated effects of Cesarean section on gene expression profiles of children in the first year of life using a generalized additive mixed model to account for the longitudinal data structure. To investigate the effect of islet autoimmunity, we compared gene expression differences between children after initiation of islet autoimmunity and age-matched children who did not develop islet autoantibodies. Finally, we compared both results to identify common regulatory patterns of Cesarean section and islet autoimmunity at the gene expression level.\n\nResultsWe identified two differentially expressed pathways in children born by Cesarean section: the pentose phosphate pathway and pyrimidine metabolism, both involved in nucleotide synthesis and cell proliferation. Islet autoantibody analysis revealed multiple differentially expressed pathways generally involved in immune processes, including both of the above-mentioned nucleotide synthesis pathways. Comparison of global gene expression signatures showed that transcriptomic changes were systematically and significantly correlated between Cesarean section and islet autoimmunity. In addition, signatures of both Cesarean section and islet autoimmunity correlated with transcriptional changes observed during activation of isolated CD4+ T lymphocytes.\n\nConclusionsWe identified coherent gene expression signatures for Cesarean section, an early risk factor for type 1 diabetes, and islet autoantibodies positivity, an obligatory stage of autoimmune response prior to the development of type 1 diabetes. Both transcriptional signatures were correlated with changes in gene expression during the activation of CD4+ T lymphocytes, reflecting common molecular changes in immune cell activation.

genomics

Network based conditional genome wide association analysis of human metabolomics

BackgroundGenome-wide association studies (GWAS) have identified hundreds of loci influencing complex human traits, however, their biological mechanism of action remains mostly unknown. Recent accumulation of functional genomics ( omics) including metabolomics data opens up opportunities to provide a new insight into the functional role of specific changes in the genome. Functional genomic data are characterized by high dimensionality, presence of (strong) statistical dependencies between traits, and, potentially, complex genetic control. Therefore, analysis of such data asks for development of specific statistical genetic methods.\n\nResultsWe propose a network-based, conditional approach to evaluate the impact of genetic variants on omics phenotypes (conditional GWAS, cGWAS). For each trait of interest, based on biological network, we select a set of other traits to be used as covariates in GWAS. The network could be reconstructed either from biological pathway databases or directly from the data. We evaluated our approach using data from a population-based KORA study (n=1,784, 1.7 M SNPs) with measured metabolomics data (151 metabolites) and demonstrated that our approach allows for identification of up to five additional loci not detected by conventional GWAS. We show that this gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.\n\nConclusionsWe demonstrate that in context of metabolomics network-based, conditional genome-wide association analysis is able to dramatically increase power of identification of loci with specific contra-intuitive pleiotropic architecture. Our method has modest computational costs, can utilize summary level GWAS data, and is applicable to other omics data types. We anticipate that application of our method to new and existing data sets will facilitate progress in understanding genetic bases of control of molecular and complex phenotypes.\n\nShort abstractWe propose a network-based, conditional approach for genome-wide analysis of multivariate omics phenotypes. Our methods can incorporate prior biological knowledge about biological pathways from external sources. We evaluated our approach using metabolomics data and demonstrated that our approach has bigger power and allows for identification of additional loci. We show that gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.

genomics