bioRxiv ScienceSearch

bioRxiv · 10.1101/500660

Analysis of >30,000 abstracts suggests higher false discovery rates for oncology journals, especially those with low impact factors

Abstract

BackgroundRecently, there has been increasing concern about the replicability, or lack thereof, of published research. An especially high rate of false discoveries has been reported in some areas motivating the creation of resource-intensive collaborations to estimate the replication rate of published research by repeating a large number of studies. The substantial amount of resources required by these replication projects limits the number of studies that can be repeated and consequently the generalizability of the findings.\n\nMethods and findingsIn 2013, Jager and Leek developed a method to estimate the empirical false discovery rate from journal abstracts and applied their method to five high profile journals. Here, we use the relative efficiency of Jager and Leeks method to gather p-values from over 30,000 abstracts and to subsequently estimate the false discovery rate for 94 journals over a five-year time span. We model the empirical false discovery rate by journal subject area (cancer or general medicine), impact factor, and Open Access status. We find that the empirical false discovery rate is higher for cancer vs. general medicine journals (p = 5.14E-6). Within cancer journals, we find that this relationship is further modified by journal impact factor where a lower journal impact factor is associated with a higher empirical false discovery rates (p = 0.012, 95% CI: -0.010, -0.001). We find no significant differences, on average, in the false discovery rate for Open Access vs closed access journals (p = 0.256, 95% CI: -0.014, 0.051).\n\nConclusionsWe find evidence of a higher false discovery rate in cancer journals compared to general medicine journals, especially those with a lower journal impact factor. For cancer journals, a lower journal impact factor of one point is associated with a 0.006 increase in the empirical false discovery rate, on average. For a false discovery rate of 0.05, this would result in over a 10% increase to 0.056. Conversely, we find no significant evidence of a higher false discovery rate, on average, for Open Access vs. closed access journals from InCites. Our results provide identify areas of research that may need of additional scrutiny and support to facilitate replicable science. Given our publicly available R code and data, others can complete a broad assessment of the empirical false discovery rate across other subject areas and characteristics of published research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hall, L. M., Hendricks, A. E.. 2018-12-29. Analysis of >30,000 abstracts suggests higher false discovery rates for oncology journals, especially those with low impact factors. https://doi.org/10.1101/500660

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

To Tweet or Not to Tweet, That is the Question: A Randomized Trial of Twitter Effects on Article Engagement in Medical Education

Many medical education journals use Twitter to garner attention for their articles. The purpose of this study was to test the effects of tweeting on article page views and downloads.\n\nThe authors conducted a randomized trial using Academic Medicine articles published in 2015. Beginning in February through May 2018, one article per day was randomly assigned to a Twitter (case) or control group. Daily, an individual tweet was generated for each article in the Twitter group that included the title, #MedEd, and a link to the article. The link delivered users to the articles landing page, which included immediate access to the HTML full text and a PDF link. The authors extracted HTML page views and PDF downloads from the publisher. To assess differences in page views and downloads between cases and controls, a time-centered approach was used, with outcomes measured at 1, 7, and 30 days.\n\nIn total, 189 articles (94 cases, 95 controls) were analyzed. After days 1 and 7, there were no statistically significant differences between cases and controls on any metric. On day 30, HTML page views exhibited a 63% increase for cases (M=14.72, SD=63.68) when compared to controls (M=9.01, SD=14.34; incident rate ratio=1.63, p=0.01). There were no differences between cases and controls for PDF downloads on day 30.\n\nContrary to the authors hypothesis, only one statistically significant difference in page views between the Twitter and control groups was found. These findings provide preliminary evidence that after 30 days a tweet can have a small positive effect on article page views.

scientific communication and education

Antibiotics and food in the American press

The emergence of antimicrobial resistant infections from food is well documented in the scientific literature but, in this kind of matter, the public opinion is an important policy driver and is vastly forged by traditional media. Here, we propose a text mining study through about 500 articles from two reference daily U.S. newspapers to assess the media coverage of this issue. Our results indicate that, since the middle of the 80s, the two journals considered here adopted a very different narrative around the issue, echoing civil society concerns in one case and the official discourse in the other.

scientific communication and education