bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.07.17.549436

Is biomedical research self-correcting? Modelling insights on the persistence of spurious science

Abstract

The reality that volumes of published research are not reproducible has been increasingly recognised in recent years, notably in biomedical science. In many fields, spurious results are common, reducing trustworthiness of reported results. While this increases research waste, a common response is that science is ultimately self-correcting, and trustworthy science will eventually triumph. While this is likely true from a philosophy of science perspective, it does not yield information on how much effort is required to nullify suspect findings, nor factors that shape how quickly science may be correcting in the publish-or-perish environment scientists operate. There is also a paucity of information on how perverse incentives of the publishing ecosystem, which reward novel positive findings over null results, shaping the ability of published science to self-correct. Precisely what factors shape self-correction of science remain obscure, limiting our ability to mitigate harms. This modelling study illuminates these questions, introducing a simple model to capture dynamics of the publication ecosystem, exploring factors influencing research waste, trustworthiness, corrective effort, and time to correction. Results from this work indicate that research waste and corrective effort are highly dependent on field-specific false positive rates and the time delay before corrective results to spurious findings are propagated. The model also suggests conditions under which biomedical science is self-correcting, and those under which publication of correctives alone cannot stem the propagation of untrustworthy results. Finally, this work models a variety of potential mitigation strategies, including researcher and publication driven interventions. Significance statementIn biomedical science, there is increasing recognition that many results fail to replicate, impeding both scientific advances and trust in science. While science is self-correcting over long time-scales, there has been little work done on the factors that shape time to correction, the scale of corrective efforts, and the research waste generated in these endeavours. Similarly, there has been little work done on quantifying factors that might reduce negative impacts of spurious science. This work takes a modeling approach to illuminate these questions, uncovering new strategies for mitigating the impact of untrustworthy research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Grimes, D. R.. 2023-07-18. Is biomedical research self-correcting? Modelling insights on the persistence of spurious science. https://doi.org/10.1101/2023.07.17.549436

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

Exploring The Association of Student Perceptions of Their Teachers' Science Instruction and Emotional Engagement and The Variance between Grade Level

AbstractPrevious research has consistently shown students have higher learning interests attitudes, and motivation toward science in elementary compared to middle and high school. However, the key predictors behind the observed differences in engagement levels across grade levels remain unclear. This study aims to identify the association between students perceptions of science instruction and their emotional engagement, as well as examine the differences across various educational stages. A multilevel random effect model was employed to investigate this association. In addition, the analysis examined whether the associations vary between grade levels. A sample of 6465 students from 25 schools participated in the study. This study shows that students in higher grades have significantly lower emotional engagement compared to third-grade students. Similarly, students in higher grades have significantly lower values on perceptions of science instruction of interesting science and understandable science decrease compared to third grade. Findings reveal significant associations between students perceptions of science instruction, both in perceiving interesting science instruction and understandable science instruction, and their emotional engagement in science learning. The interaction term (Perception of Interesting Science Instruction x Grade) of Grades four, six through eight, and grade 12 are statistically significantly related to emotional engagement, indicating that more positive slopes correspond to high-grade level than third-grade students. The increase in students perception of interesting science has the steepest slope of improvement in students emotional engagement in 6th grade. We find no statistically significant difference between grades in students perception of understandable science and emotional engagement. Interesting and understandable science classes have positive correlations with students emotional engagement in science learning. For future studies, it is worth exploring science instruction during transition grades to better address and mitigate the observed difference in emotional engagement.

scientific communication and education↗

Retrospective Analysis of the Effects of BWF Interdisciplinary Postdoctoral to Faculty Transition Awards on Future Funding Success

Established by the Burroughs Wellcome Fund (BWF) in 2001, the Career Award at the Scientific Interface (CASI) is a career development award for scientists with doctoral training in the physical/mathematical/computational sciences or engineering conducting postdoctoral research in the biological sciences. The goal of the program is to support early career scientists interested in pursuing an independent research career with an interdisciplinary focus. In order to assess the benefit of the CASI award on recipients, the authors undertook a retrospective analysis of the funding data for CASI recipients to evaluate success against matching cohorts. These cohorts included applicants who succeeded to the final interview stage but were ultimately unsuccessful (interviewed), applicants who submitted proposals but did not make it to the final interview stage (proposal declined), and a randomly selected dataset of researchers from a comparable program, the highly competitive Pathway to Independence Award (K99/R00) from the National Institutes of Health (NIH). The results indicate that CASI recipients outperformed unsuccessful applicants and their K99/R00 counterparts in federal grant rates and overall grant dollars. The authors conclusion affirms that the CASI mechanism and BWF support successfully achieve the objective of invigorating the careers of young investigators, resulting in tangible downstream long-term effects.

scientific communication and education↗