bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.02.01.578440

Structured Peer Review: Pilot results from 23 Elsevier Journals

Abstract

BackgroundReviewers rarely comment on the same aspects of a manuscript, making it difficult to properly assess manuscripts quality and the quality of the peer review process. It was the goal of this pilot study to evaluate structured peer review implementation by: 1) exploring if and how reviewers answered structured peer review questions, 2) analysing reviewer agreement, 3) comparing that agreement to agreement before implementation of structured peer review, and 4) further enhancing the piloted set of structured peer review questions. MethodsStructured peer review consisting of 9 questions was piloted in August 2022 in 220 Elsevier journals. We randomly selected 10% of these journals across all fields and IF quartiles and included manuscripts that in the first 2 months of the pilot received 2 reviewer reports, leaving us with 107 manuscripts belonging to 23 journals. Eight questions had open ended fields, while the ninth question (on language editing) had only a yes/no option. Reviews could also leave Comments-to-Author and Comments-to-Editor. Answers were qualitatively analysed by two raters independently. ResultsAlmost all reviewers (n=196, 92%) filled out the answers to all questions even though these questions were not mandatory in the system. The longest answer (Md 27 words, IQR 11 to 68) was for reporting methods with sufficient details for replicability or reproducibility. Reviewers had highest (partial) agreement (of 72%) for assessing the flow and structure of the manuscript, and lowest (of 53%) for assessing if interpretation of results are supported by data, and for assessing if statistical analyses were appropriate and reported in sufficient detail (also 52%). Two thirds of reviewers (n=145, 68%) filled out the Comments-to-Author section, of which 105 (49%) resembled traditional peer review reports. Such reports contained a Md of 4 (IQR 3 to 5) topics covered by the structured questions. Absolute agreement regarding final recommendations (exact match of recommendation choice) was 41%, which was higher than what those journals had in the period of 2019 to 2021 (31% agreement, P=0.0275). ConclusionsOur preliminary results indicate that reviewers adapted to the new format of review successfully, and answered more topics than they covered in their traditional reports. Individual question analysis indicated highest disagreement regarding interpretation of results and conducting and reporting of statistical analyses. While structured peer review did lead to improvement in reviewer final recommendation agreements, this was not a randomized trial, and further studies should be done to corroborate this. Further research is also needed to determine if structured peer review leads to greater knowledge transfer or better improvement of manuscripts.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Malicki, M., Mehmani, B.. 2024-02-04. Structured Peer Review: Pilot results from 23 Elsevier Journals. https://doi.org/10.1101/2024.02.01.578440

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

An All-In-One Software Solution for Automated Processing of LA-ICP-TOF-MS datasets

LA-ICP-TOF-MS provides rapid, high resolution elemental analysis of biological and non-biological samples. However, accurate real-time data analysis frequently requires the user to account for several instrumental and experimental variables that can change during data acquisition. AutoSpect is a novel software tool designed to automate the processing and fitting of LA-ICP-TOF-MS data, addressing key challenges such as time-dependent spectral drift, instrument sensitivity drift calibration inaccuracies, and peak deconvolution, enabling researchers to rapidly and accurately process complex datasets. The tool is optimized to be robustly applicable across scientific fields (e.g., geochemistry, biology, and materials science), providing a streamlined solution for end users seeking to maximize the potential of LA-ICP-TOF-MS for high-resolution elemental mapping and isotopic analysis. Significance to JAASAnalysis of fast transient signals using laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) has become mainstream for elemental mapping. Advancements in LA-ICP-TOF-MS technology continue to accelerate the collective understanding of the role inorganic chemistry plays in dynamic processes. To ensure accurate quantitative results, the vast amount of complex spectral data generated requires elegant solutions to perform a variety of functions including data partitioning, peak fitting, drift correction, mass-to-charge calibration, peak profiling, and spectral fitting. AutoSpect is an all-in-one software solution that provides high level automation with a user-friendly graphical interface to perform complex data analyses for ICP-TOF-MS datasets.

scientific communication and education↗

Could instructor talk drive CURE effectiveness? A comparative study of instructor talk in introductory lab courses

Course-based undergraduate research experiences (CUREs) are thought to enhance students motivation to continue in college, in science, and in research. Yet, how CUREs enhance student motivation is largely undefined. Theories of instructor immediacy, self-efficacy, and task values suggest that CURE instructors may talk in ways that influence students motivational beliefs. We characterized the non-content related talk of instructors teaching 48 introductory biology lab courses, half CUREs and half non-CUREs. We identified 14 types of instructor talk that fit these theoretical perspectives: fostering students closeness with their instructor (i.e., immediacy talk), building students confidence in their scientific abilities (i.e., self-efficacy talk), and promoting students sense of worth in their work (i.e., task value talk). Course type had a medium effect on talk type, with CURE instructors utilizing more immediacy, self-efficacy, and task values talk than non-CURE instructors but also showing more variation in these types of talk. Our results suggest that motivation-related instructor talk is more prevalent in CUREs than non-CUREs, but wide variation in CURE instructor talk indicates additional investigation is needed before non-content talk can be considered a mechanism for the motivational influences of CUREs. HIGHLIGHTThis study compares non-content instructor talk in CURE and non-CURE lab courses using immediacy, self-efficacy, and task value theories. CURE instructors use more talk than non-CURE instructors, but variation in CURE instructor talk leaves open the question of whether talk is a causal factor in the motivational influence of CURE instruction.

scientific communication and education↗