bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.09.03.556128

How prior and p-value heuristics are used when interpreting data

Abstract

Scientific conclusions are based on the ways that researchers interpret data, a process that is shaped by psychological and cultural factors. When researchers use shortcuts known as heuristics to interpret data, it can sometimes lead to errors. To test the use of heuristics, we surveyed 623 researchers in biology and asked them to interpret scatterplots that showed ambiguous relationships, altering only the labels on the graphs. Our manipulations tested the use of two heuristics based on major statistical frameworks: (1) the strong prior heuristic, where a relationship is viewed as stronger if it is expected a priori, following Bayesian statistics, and (2) the p-value heuristic, where a relationship is viewed as stronger if it is associated with a small p-value, following null hypothesis statistical testing. Our results show that both the strong prior and p-value heuristics are common. Surprisingly, the strong prior heuristic was more prevalent among inexperienced researchers, whereas its effect was diminished among the most experienced biologists in our survey. By contrast, we find that p-values cause researchers at all levels to report that an ambiguous graph shows a strong result. Together, these results suggest that experience in the sciences may diminish a researchers Bayesian intuitions, while reinforcing the use of p-values as a shortcut for effect size. Reform to data science training in STEM could help reduce researchers reliance on error-prone heuristics. Significance StatementScientific researchers must interpret data and statistical tests to draw conclusions. When researchers use shortcuts known as heuristics, it can sometimes lead to errors. To test how this occurs, we asked biologists to interpret graphs that showed an ambiguous relationship between two variables, and report whether the relationship was strong, weak, or absent. We altered features of the graph to test whether prior expectations or a statistic called the p-value could influence their interpretations. Our results indicate that both prior expectations and p-values can increase the probability that researchers will report that ambiguous data shows a strong result. These findings suggest that current training and research practices promote the use of error-prone shortcuts in decision-making.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hermer, E., Irwin, A. A., Roche, D. G., Dakin, R.. 2023-09-06. How prior and p-value heuristics are used when interpreting data. https://doi.org/10.1101/2023.09.03.556128

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

Trends in hoaxes of academic communication

Academic journals use peer review to weed out false information, but peer review and other editorial processes are normally confidential. Therefore, individuals sometimes create hoaxes to test whether editorial processes are as robust as they are claimed, or whether they are done at all. This article tracks the occurrence of hoaxes aimed at scholarly publishers and academic conferences since 2000. Since 2009, successful hoaxes usually appeared at a year of one or more a year, usually motivated by academics or journalists exposing so-called "predatory" journals. The apparent rise in the number of hoaxes reflects a lack of transparency in editorial processes at both legitimate and "predatory" journals. Reaction of academic communities to hoaxes varies widely depending on the perceived intent of the target of the hoaxes and whether the hoax demonstrates what the hoaxer claims.

scientific communication and education↗

Capacity building needed to reap the benefits of access to biodiversity collections

SummaryO_LIThis research examines biodiversity specimens from two areas of the Caribbean to understand patterns of collection and the roles of the people involved. Using open data from the Global Biodiversity Information Facility (GBIF) and Wikidata, we aimed to uncover geographic and historical trends in specimen use. This study aims to provide concrete evidence to guide collaboration between collection-holding institutions and the communities that need their resources most. C_LIO_LIWe analysed biodiversity specimens from Montserrat and the Cayman Islands in three steps. First, we extracted specimen data from GBIF, disambiguated collector names, and linked them to unique biographical entries. Next, we connected collectors to their publications and specimens. Finally, we analysed the modern use of these specimens through citation data, mapping author affiliations and research themes. C_LIO_LISpecimens are predominantly housed in the Global North and were initially used by their collectors, whose focus was largely on taxonomy and biogeography. With digitisation, use of these collections remains concentrated in the Global North and covers a broader range of subjects, although Brazil and China stand out as significant users of digital collection data compared to other similar countries. C_LIO_LIThe availability of open digital data from collections in the Global North has led to a substantial increase in the reuse of these data across biodiversity science. Nonetheless, most research using these data is still conducted in the Global North. For the non-monetary benefits of digitisation to extend to the countries of origin, capacity building in the Global South is crucial, Open Data alone are insufficient. C_LI Societal Impact StatementDigital biodiversity data from herbaria and museums hold significant potential for nature conservation in the Global South, yet many regions, like Montserrat and the Cayman Islands in the Caribbean, are, for multiple reasons, unable to fully leverage this information. This lack of skills and resources limits local conservation efforts, showing the need for more investment in training, facilities, and expertise. Although past funding has helped improve coordination and build skills, our findings show that more work is needed to make sure conservation in these biodiverse areas can continue in the long term.

scientific communication and education↗