bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.04.19.590240

Enabling preprint discovery, evaluation, and analysis with Europe PMC

Abstract

Preprints provide an indispensable tool for rapid and open communication of early research findings. Preprints can also be revised and improved based on scientific commentary uncoupled from journal-organised peer review. The uptake of preprints in the life sciences has increased significantly in recent years, especially during the COVID-19 pandemic, when immediate access to research findings became crucial to address the global health emergency. With ongoing expansion of new preprint servers, improving discoverability of preprints is a necessary step to facilitate wider sharing of the science reported in preprints. To address the challenges of preprint visibility and reuse, Europe PMC, an open database of life science literature, began indexing preprint abstracts and metadata from several platforms in July 2018. Since then, Europe PMC has continued to increase coverage through addition of new servers, and expanded its preprint initiative to include the full text of preprints related to COVID-19 in July 2020 and then the full text of preprints supported by the Europe PMC funder consortium in April 2022. The preprint collection can be searched via the website and programmatically, with abstracts and the open access full text of COVID-19 and Europe PMC funder preprint subsets available for bulk download in a standard machine-readable JATS XML format. This enables automated information extraction for large-scale analyses of the preprint corpus, accelerating scientific research of the preprint literature itself. This publication describes steps taken to build trust, improve discoverability, and support reuse of life science preprints in Europe PMC. Here we discuss the benefits of indexing preprints alongside peer-reviewed publications, and challenges associated with this process.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Levchenko, M., Parkin, M., McEntyre, J., Harrison, M.. 2024-04-19. Enabling preprint discovery, evaluation, and analysis with Europe PMC. https://doi.org/10.1101/2024.04.19.590240

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

An All-In-One Software Solution for Automated Processing of LA-ICP-TOF-MS datasets

LA-ICP-TOF-MS provides rapid, high resolution elemental analysis of biological and non-biological samples. However, accurate real-time data analysis frequently requires the user to account for several instrumental and experimental variables that can change during data acquisition. AutoSpect is a novel software tool designed to automate the processing and fitting of LA-ICP-TOF-MS data, addressing key challenges such as time-dependent spectral drift, instrument sensitivity drift calibration inaccuracies, and peak deconvolution, enabling researchers to rapidly and accurately process complex datasets. The tool is optimized to be robustly applicable across scientific fields (e.g., geochemistry, biology, and materials science), providing a streamlined solution for end users seeking to maximize the potential of LA-ICP-TOF-MS for high-resolution elemental mapping and isotopic analysis. Significance to JAASAnalysis of fast transient signals using laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) has become mainstream for elemental mapping. Advancements in LA-ICP-TOF-MS technology continue to accelerate the collective understanding of the role inorganic chemistry plays in dynamic processes. To ensure accurate quantitative results, the vast amount of complex spectral data generated requires elegant solutions to perform a variety of functions including data partitioning, peak fitting, drift correction, mass-to-charge calibration, peak profiling, and spectral fitting. AutoSpect is an all-in-one software solution that provides high level automation with a user-friendly graphical interface to perform complex data analyses for ICP-TOF-MS datasets.

scientific communication and education↗

Could instructor talk drive CURE effectiveness? A comparative study of instructor talk in introductory lab courses

Course-based undergraduate research experiences (CUREs) are thought to enhance students motivation to continue in college, in science, and in research. Yet, how CUREs enhance student motivation is largely undefined. Theories of instructor immediacy, self-efficacy, and task values suggest that CURE instructors may talk in ways that influence students motivational beliefs. We characterized the non-content related talk of instructors teaching 48 introductory biology lab courses, half CUREs and half non-CUREs. We identified 14 types of instructor talk that fit these theoretical perspectives: fostering students closeness with their instructor (i.e., immediacy talk), building students confidence in their scientific abilities (i.e., self-efficacy talk), and promoting students sense of worth in their work (i.e., task value talk). Course type had a medium effect on talk type, with CURE instructors utilizing more immediacy, self-efficacy, and task values talk than non-CURE instructors but also showing more variation in these types of talk. Our results suggest that motivation-related instructor talk is more prevalent in CUREs than non-CUREs, but wide variation in CURE instructor talk indicates additional investigation is needed before non-content talk can be considered a mechanism for the motivational influences of CUREs. HIGHLIGHTThis study compares non-content instructor talk in CURE and non-CURE lab courses using immediacy, self-efficacy, and task value theories. CURE instructors use more talk than non-CURE instructors, but variation in CURE instructor talk leaves open the question of whether talk is a causal factor in the motivational influence of CURE instruction.

scientific communication and education↗