bioRxiv Science⌕ Search

Biology subjects

Duperier, C.

Publications and source records attributed to Duperier, C..

2 recordsLinked to original sources

Suggesting disease associations for overlooked metabolites using literature from metabolic neighbours

In human health research, metabolic signatures extracted from metabolomics data are a strong-added value for stratifying patients and identifying biomarkers. Nevertheless, one of the main challenges is to interpret and relate these lists of discriminant metabolites to pathological mechanisms. This task requires experts to combine their knowledge with information extracted from databases and the scientific literature. However, we show that a large fraction of metabolites are rarely or never mentioned in the literature. Consequently, these overlooked metabolites are often set aside and the interpretation of metabolic signatures is restricted to a subset of the significant metabolites. To suggest potential pathological phenotypes related to these understudied metabolites, we extend the guilt by association principle to literature information by using a Bayesian framework. With this approach, we suggest more than 35,000 associations between 1,047 overlooked metabolites and 3,288 diseases (or disease families). All these newly inferred associations are freely available on the FORUM ftp server (See information at https://github.com/eMetaboHUB/Forum-LiteraturePropagation.).

bioinformatics↗

FORUM: Building a Knowledge Graph from public databases and scientific literature to extract associations between chemicals and diseases

Metabolomics studies aim at reporting a metabolic signature (list of metabolites) related to a particular experimental condition. These signatures are instrumental in the identification of biomarkers or classification of individuals, however their biological and physiological interpretation remains a challenge. Overcoming this challenge is critical when aiming to associate metabolic signatures with potential pathological outcomes. To support this task, we introduce FORUM: a Knowledge Graph (KG) providing a semantic representation of relations between chemicals and biomedical concepts, built from a federation of life science databases and scientific literature repositories. An important number of scientific articles discuss relations between chemical compounds and biomedical concepts in various contexts, from biomarkers to therapeutic uses. The extraction of these statements and their interconnection in a graph structure can thus allow us to identify and explore relations strongly supported in the scientific literature. The use of a Semantic Web framework on biological data allows us to apply ontological based reasoning to infer new relations between entities. We show that these new relations provide different levels of abstraction and could open the path to new hypotheses. We estimate the statistical relevance of each extracted relation, explicit or inferred, using an enrichment analysis, and instantiate them as new knowledge in the KG to support results interpretation/further inquiries. Beyond this result, FORUM can also provide insights into complex biological questions and the extracted information could then be used for further developments. Containing more than 8 billion triples and providing more than 8 million relations, FORUM leverages the increasing availability of linked datasets in life science and is built in agreement with FAIR principles. A web interface to browse and download the extracted relations, as well as a SPARQL endpoint to directly probe the whole FORUM knowledge graph, are available at https://forum-webapp.semantic-metabolomics.fr. The code needed to reproduce the triplestore is available at https://github.com/eMetaboHUB/Forum-DiseasesChem.

bioinformatics↗