bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.10.16.618663

Creating a biomedical knowledge base by addressing GPT inaccurate responses and benchmarking context

Abstract

We created GNQA, a generative pre-trained transformer (GPT) knowledge base driven by a performant retrieval augmented generation (RAG) with a focus on aging, dementia, Alzheimers and diabetes. We uploaded a corpus of three thousand peer reviewed publications on these topics into the RAG. To address concerns about inaccurate responses and GPT hallucinations, we implemented a context provenance tracking mechanism that enables researchers to validate responses against the original material and to get references to the original papers. To assess the effectiveness of contextual information we collected evaluations and feedback from both domain expert users and citizen scientists on the relevance of GPT responses. A key innovation of our study is automated evaluation by way of a RAG assessment system (RAGAS). RAGAS combines human expert assessment with AI-driven evaluation to measure the effectiveness of RAG systems. When evaluating the responses to their questions, human respondents give a "thumbs-up" 76% of the time. Meanwhile, RAGAS scores 90% on answer relevance on questions posed by experts. And when GPT-generates questions, RAGAS scores 74% on answer relevance. With RAGAS we created a benchmark that can be used to continuously assess the performance of our knowledge base. Full GNQA functionality is embedded in the free GeneNetwork.org web service, an open-source system containing over 25 years of experimental data on model organisms and human. The code developed for this study is published under a free and open-source software license at https://git.genenetwork.org/gn-ai/tree/README.md.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Darnell, S. S., Prins, J. P., Suh, E., Huang, P., Williams, R. W., Overall, R., Chen, H., Garrison, E., Guarracino, A., Villani, F., Muli, P., Ashbrook, D. G., Colonna, V., Batten, C., Sen, S., Muriithi, F. M., Yousefi, S., Nijveen, H., Lisso, F., Isaac, A., Kabui, A., Kilungi, M. B., Kibet, A., Umar, M., Muhia, B.. 2024-10-18. Creating a biomedical knowledge base by addressing GPT inaccurate responses and benchmarking context. https://doi.org/10.1101/2024.10.16.618663

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

The Research Mind: A Multi-Dimensional Framework for Evaluating Research Quality, Productivity, and Integrity by Mitigating Citation Bias

The top 2% researchers list from Professor John P. A. Ioannidis from Stanford University captures a lot of attention worldwide. Professor Ioannidis introduced a new marker to evaluate a researchers contributions named the composite score (C-Score). The C-Score provided an excellent metric to determine where an author stands in comparison to others within the global research community. However, while the C-Score effectively ranks researchers, it is not particularly helpful for new researchers seeking to evaluate potential mentors or guides. Emerging researchers want to know if potential mentors prioritize quantity or quality. They also seek clarity about research output requirements when joining a team. ObjectiveWe address this gap with The Research Mind framework. It introduces two new metrics: the S-Score (measures research quantity and productivity) and the Q-Score (measures research quality and impact). These metrices help new researchers choose mentors wisely by going beyond traditional ranking systems. MethodsWe developed the S-Score to measure maximum average first-author publications over any three-year period. This provides insights into productivity expectations. We also created the Q-Score to evaluate maximum median citations for first/last author works over three years. This indicates research quality and impact. Additionally, we implemented comprehensive self-citation analysis to assess research integrity. The framework was built using OpenAlex database and includes network analysis capabilities for collaboration pattern visualization. ResultsOur analysis uncovered significant differences between C-Score rankings and a researchers suitability for mentorship. We found that some researchers had high C-Scores but also had S-Scores of 10, meaning they published 10 papers per year. This is an unusually high productivity rate that may be unsustainable for most researchers. In contrast, researchers publishing one paper per year had Q-Scores above 50, making them potentially better mentors for quality-focused research. Our self-citation analysis revealed concerning patterns, with high self-citation rates (>20%) indicating potentially problematic research practices. ConclusionThe Research Mind framework provides essential complementary metrics to Ioannidiss C-Score system, enabling new researchers to evaluate potential guides/mentors based on productivity expectations (S-Score) and quality focus (Q-Score) rather than solely on composite rankings. This approach will help emerging researchers identify mentors whose research philosophy and output expectations align with their career goals and capabilities. Hence, The Research Mind framework addresses a critical gap in academic mentorship selection. To facilitate widespread adoption and accessibility, we have developed a comprehensive web application. This app allows researchers to easily search, analyze, and compare these metrics for any author/researcher. The platform is freely available at www.theresearchmind.com/trm-app, providing an intuitive interface for exploring S-Scores, Q-Scores, self-citation patterns, collaboration networks, and overall evaluation metrics to support informed academic decision-making. PVLDB Reference FormatSanjay Rathee and Chanchal Chaudhary. The Research Mind: A Multi-Dimensional Framework for Evaluating Research Quality, Productivity, and Integrity by Mitigating citation bias. PVLDB, 14(1): XXX-XXX, 2020. doi:XX.XX/XXX.XX PVLDB Artifact AvailabilityThe source code, data, and/or other artifacts have been made available at URL_TO_YOUR_ARTIFACTS.

scientific communication and education↗

Researchers' perspectives on preregistration in animal research

Preregistration is arguably one of the most promising and impactful Open Science practices. Defined as the a priori registration of study designs and analysis plans, preregistration has long been established as standard practice in clinical human research and is increasingly taken up in other fields of science. Despite growing evidence suggesting that preregistration can mitigate questionable research practices, and the existence of two platforms targeting animal studies, preregistration remains uncommon in animal research. In light of the reproducibility crisis and calls for more transparency and rigor also in animal research, preregistration represents a potentially promising step forward. However, implementing such policies without understanding their impact carries potential risks. It is, therefore, essential to uncover the strengths, weaknesses, opportunities, and threats of preregistration before advancing its implementation in animal research. This current study addressed this need as part of a larger feasibility project on preregistration of animal experiments in Switzerland, and aimed to: 1) assess the researchers experiences with preregistration; 2) examine their attitudes, subjective norms, perceived behavioral control, intentions, motivations, and perceived obstacles regarding preregistration; 3) explore associations between these psychosocial constructs and relevant background characteristics; 4) identify perceived facilitators and barriers to preregistration; and 5) summarize researchers suggestions for improving preregistration. A preregistered cross-sectional online survey was conducted among all registered study directors of ongoing animal experiments in Switzerland. Of the 1,385 invited study directors, 418 completed the survey (30.2% return rate; 41% female; age M = 47.1, SD = 9.52). Among them, 39.2% had never heard of preregistration, and only 10% had preregistered studies before participating in the survey. Bureaucratic burden (77.6%), time costs (71.4%), and low flexibility (65.7%) were the most common reported barriers to preregistration. On average, participants described rather unfavorable attitudes towards preregistration, negative subjective norms, relatively low perceived behavioral control, weak intention, and limited motivation to preregister, along with high perceived obstacles. Participants who had never preregistered a study, as well as those with more research experience, showed more negative scores on all assessed psychosocial constructs related to preregistration. Our findings offer guidance on promising measures to enhance acceptance of preregistration among animal researchers, including raising awareness, offering education and training, and facilitating procedures, to enhance the acceptance of preregistration among animal researchers.

scientific communication and education↗