bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.04.29.067843

A study's got to know its limitations

Abstract

Background All research has room for improvement, but authors do not always clearly acknowledge the limitations of their work. In this brief report, we sought to identify the prevalence of limitations statements in the medRxiv COVID-19 SARS-CoV-2 dataset.Methods We combined automated methods with manual review to analyse manuscripts for the presence, or absence, either of a defined limitations section in the text, or as part of the general discussion.Results We identified a structured limitations statement in 28% of the manuscripts, and overall 52% contained at least one mention of a study limitation. Over one-third of manuscripts contained none of the terms that might typically be associated with reporting of limitations. Overall our method performed with precision of 0.97 and recall of 0.91.Conclusion The presence or absence of limitations statements can be identified with reasonable confidence using automated tools. We suggest that it might be beneficial to require a defined, structured statement about study limitations, either as part of the submission process, or clearly delineated within the manuscript.Competing Interest StatementThe authors are the founders and shareholders of Scholarcy Limited. Scholarcy technology was used to extract the information from the manuscripts used in this study.View Full Text

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gooch, P., Warren-Jones, E.. 2020-05-02. A study's got to know its limitations. https://doi.org/10.1101/2020.04.29.067843

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

The data-index: an author-level metric that values impactful data and incentivises data sharing

Author-level metrics are a widely used measure of scientific success. The h-index, and its variants, measure publication output (number of publications) and impact (number of citations), and these are often used to allocate funding or jobs. Here we argue that the emphasis on publication output and impact hinders progress in the fields of ecology and evolution as it disincentivises two fundamental practices: generating long-term datasets and sharing data. We describe a new author-level metric, the data-index, which values dataset output and impact and promotes generating and sharing data as a result. It is designed to complement other metrics of scientific success, as scientific contributions are diverse and our value system should reflect that. Future work should focus on designing alternative metrics that value our wider merits, such as communicating our research, informing policy, mentoring other scientists, and providing open-access code and tools.

scientific communication and education

A laboratory module that explores RNA interference and codon optimization through fluorescence microscopy using Caenorhabditis elegans

Scientific research experiences are beneficial to students allowing them to gain laboratory and problem-solving skills, as well as foundational research skills in a team-based setting. We designed a laboratory module to provide a guided research experience to stimulate curiosity, introduce students to experimental techniques, and provide students with foundational skills needed for higher levels of guided inquiry. In this laboratory module, students learn about RNA interference (RNAi) and codon optimization using the research organism Caenorhabditis elegans (C. elegans). Students are given the opportunity to perform a commonly used method of gene downregulation in C. elegans where they visualize gene depletion using fluorescence microscopy and quantify the efficacy of depletion using quantitative image analysis. The module presented here educates students on how to report their results and findings by generating publication quality figures and figure legends. The activities outlined exemplify ways by which students can improve their critical thinking, data interpretation, and technical skills, all of which are beneficial for future laboratory classes, independent inquiry-based research projects, and careers in the life sciences and beyond.

scientific communication and education