bioRxiv Science⌕ Search

Biology subjects

Heverin, M.

Publications and source records attributed to Heverin, M..

2 recordsLinked to original sources

Mining impactful discoveries from the biomedical literature

MotivationLiterature-Based Discovery (LBD) aims to help researchers to identify relations between concepts which are worthy of further investigation by text-mining the biomedical literature. While the LBD literature is rich and the field is considered mature, standard practice in the evaluation of LBD methods is methodologically poor and has not progressed on par with the domain. The lack of properly designed and decent-sized benchmark dataset hinders the progress of the field and its development into applications usable by biomedical experts. ResultsThis work presents a method for mining past discoveries from the biomedical literature. It leverages the impact made by a discovery, using descriptive statistics to detect surges in the prevalence of a relation across time. This method allows the collection of a large amount of time-stamped discoveries which can be used for LBD evaluation or other applications. The validity of the method is tested against a baseline representing the state of the art "time sliced" method. AvailabilityThe source data used in this article are publicly available. The implementation and the resulting data are published under open-source license: https://github.com/erwanm/medline-discoveries (code) https://zenodo.org/record/5888572 (datasets). An online exploration tool is also provided at https://brainmend.adaptcentre.ie/. Contacterwan.moreau@adaptcentre.ie

bioinformatics↗

Literature-Based Discovery beyond the ABC paradigm: a contrastive approach

Literature-Based Discovery (LBD) aims to help researchers to identify relations between concepts which are worthy of further investigation by text-mining the biomedical literature. The vast majority of the LBD research follows the ABC model: a relation (A,C) is a candidate for discovery if there is some intermediate concept B which is related to both A and C. The ABC model has been successful in applications where the search space is strongly constrained, but there is limited evidence about its usefulness when applied in a broader context. Through a case study of 8 recent discoveries related to neurodegenerative diseases (NDs), we show the limitations of the ABC model in an open-ended context. The study emphasizes the impact of the choice of source data and extraction method on the resulting knowledge base: different "views" of the biomedical literature offer different levels of accuracy and coverage. We propose a novel contrastive approach which leverages these differences between "views" in order to target relations between concepts of interest. We explore various parameters and demonstrate the relevance of our approach through quantitative evaluation on the 8 target discoveries. The source data used in this article are publicly available. The different parts of the software used to process the data are published under open-source license and provided with detailed instructions. The main code for this paper is available at https://github.com/erwanm/lbd-contrast (required dependencies are detailed in the documentation). A prototype of the system is also provided as an online exploration tool at brainmend.adaptcentre.ie.

bioinformatics↗