bioRxiv Science⌕ Search

bioRxiv · 10.1101/2021.01.19.427362

Alignment of biomedical data repositories with open, FAIR, citable and trustworthy principles

Abstract

Increasing attention is being paid to the operation of biomedical data repositories in light of efforts to improve how scientific data is handled and made available for the long term. Multiple groups have produced recommendations for functions that biomedical repositories should support, with many using requirements of the FAIR data principles as guidelines. However, FAIR is but one set of principles that has arisen out of the open science community. They are joined by principles governing open science, data citation and trustworthiness, all of which are important aspects for biomedical data repositories to support. Together, these define a framework for data repositories that we call OFCT: Open, FAIR, Citable and Trustworthy. Here we developed an instrument using the open source PolicyModels toolkit that attempts to operationalize key aspects of OFCT principles and piloted the instrument by evaluating eight biomedical community repositories listed by the NIDDK Information Network (dkNET.org). Repositories included both specialist repositories that focused on a particular data type or domain, in this case diabetes and metabolomics, and generalist repositories that accept all data types and domains. The goal of this work was both to obtain a sense of how much the design of current biomedical data repositories align with these principles and to augment the dkNET listing with additional information that may be important to investigators trying to choose a repository, e.g., does the repository fully support data citation? The evaluation was performed from March to November 2020 through inspection of documentation and interaction with the sites by the authors. Overall, although there was little explicit acknowledgement of any of the OFCT principles in our sample, the majority of repositories provided at least some support for their tenets.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Murphy, F., Bar-Sinai, M., Martone, M. E.. 2021-01-20. Alignment of biomedical data repositories with open, FAIR, citable and trustworthy principles. https://doi.org/10.1101/2021.01.19.427362

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

MolecularWebXR: Multiuser discussions about chemistry and biology in immersive and inclusive VR

MolecularWebXR is a new website for education, science communication and scientific peer discussion in chemistry and biology, based on modern web-based Virtual Reality (VR) and Augmented Reality (AR). With no installs as it is all web-served, MolecularWebXR enables multiple users to simultaneously explore, communicate and discuss concepts about chemistry and biology in immersive 3D environments, by manipulating and passing around objects with their bare hands and pointing at different elements with natural hand gestures. User may either be present in the same real space or distributed around the world, in the latter case talking naturally with each other thanks to built-in audio features. Although MolecularWebXR is most immersive when running in the web browsers of high-end AR/VR headsets, its WebXR core also allows participation by users with consumer devices such as smartphones, possibly inserted into cardboard goggles for deeper immersivity, or even in computers and tablets. MolecularWebXR comes with preset VR rooms that cover topics from general, inorganic and organic chemistry, biophysics and structural biology, and general biology; besides, new content can be added at will through moleculARwebs PDB2AR tool or by contacting the lead authors. We verified MolecularWebXRs ease of use and versatility by people aged 12-80 years old in entirely virtual sessions or in mixed real-virtual sessions at various science outreach events, in courses at the bachelor, masters and early doctoral levels, in scientific collaborations, and in conference lectures. MolecularWebXR is available for free use without registration at https://molecularwebxr.org, and a blog post version of this preprint with embedded videos is available at https://go.epfl.ch/molecularwebxr-blog-post.

scientific communication and education↗

Willingness of Japanese people in their 20s, 30s and 40s to pay for genetically modified foods

The application of genetically modified (GM) technology to food products has increased worldwide. The adaptation has extended to conventional grains and animal products, such as salmon. However, in Japan, the publics acceptance of GM foods is low and experts and policymakers need to know the publics preference for various types of GM foods. Therefore, this study aims to clarify and compare the preferences for various GM foods among Japanese people in their 20s, 30s, and 40s, using the Willingness-to-pay (WTP) indicator. An online survey with 1122 valid responses from people in their 20s-40s was used for analysis. The results showed that the percentages of willingness to purchase various items were as follows - GM blue roses (n = 628; 56%), tomatoes fertilized with GM oil cake (n = 519; 46.3%), potato chips made from GM potatoes (n = 489; 43.6%), chicken thighs fed on GM corn (n = 472; 42.1%), corn flakes made from GM corn (n = 471; 42.0%), GM apples (n = 420; 37.4%), wine brewed with GM yeast (n = 416; 37.1%), GM tomatoes (n = 408; 36.4%), GM chicken thighs (n = 360; 32.1%), GM salmon (n = 349; 31.1%). Comparing the WTP discount rates for animal and plant foods, it was around 20% for animal products (GM chicken thighs = 17.8% and GM salmon = 19.5%), and 15-35% for plant foods (GM corn flakes = 24.1%, GM tomatoes = 17.2%, GM potato chips = 16.4%, GM yeast wine = 36.7%, GM apples = 25.3%). The WTP discount rates were 11.9% for tomatoes fertilized with GM oil cake compared with 17.2% for GM tomatoes, and 21.5% for GM-fed animal products compared with 17.8% for chicken thighs fed with GM corn. Therefore, the WTP values of GM animal foods were lower than GM plant foods and ornamental products.

scientific communication and education↗