bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.04.17.719279

ReviewBench: An Extensible Framework for Benchmarking Human and AI Manuscript Review

Abstract

AO_SCPLOWBSTRACTC_SCPLOWThe volume of scientific manuscripts is rising faster than the available pool of expert reviewers, and AI tools are emerging as a possible response, ranging from frontier large language models applied directly to peer review to purpose-built multi-agent systems. Scalable, standardized benchmarks are needed to regularly evaluate how these tools compare to one another and to human reviewers. We present ReviewBench, an open-source, venue-agnostic framework that compares human and AI reviews across structure, alignment with a papers major claims, impact, and critique category. We apply ReviewBench to 145,021 review comments from human reviewers, frontier large language models (GPT-5.2, and Gemini 3 Pro), and Reviewer3.com (R3), a multi-agent peer review system. The dataset spans papers in computer science (ICLR 2025, n = 1,000), social science (Nature Human Behaviour, n = 142), and life science (eLife, n = 1,000). Across disciplines, AI reviews are more structured and engage more directly with a papers major claims, with R3 more often surfacing consequential comments, defined as comments capable of undermining those claims. When restricting to critical comments, however, human reviewers rank first on consequential rate on more individual papers than any AI source, despite a lower average. We identify a bimodal reviewer distribution with peaks near 0% and 100%, indicating that many reviewers outperform AI on this metric, but a substantial fraction of reviewers near 0% brings the average down. Critique typing demonstrates systematic differences, where humans emphasize contribution and clarity, while AI emphasizes validity, sufficiency, and transparency. Together, these findings argue against framing AI as a replacement for human review and instead support a complementary model in which AI scales technical verification of major claims while human judgment remains essential for evaluating contribution and shaping editorial decisions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Khalil, N. N., Reed, T. J., Ciccozzi, M. R.. 2026-04-20. ReviewBench: An Extensible Framework for Benchmarking Human and AI Manuscript Review. https://doi.org/10.64898/2026.04.17.719279

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Multi-Lab Testing of Early Preclinical Discoveries Identifies Promising Treatments

A fundamental challenge in drug development is the frequent failure of early laboratory research to translate into clinical benefit. One promising solution is to confirm findings from exploratory single-laboratory studies across multiple laboratories before clinical testing. We investigated this approach following the conduct of preclinical multi-laboratory studies across different fields of medicine. For this, we evaluated effect sizes, experimental rigor, and a set of criteria to identify determinants of confirmation success. When tested under increased rigor, only a fraction of multi-laboratory studies confirmed the initial results. The underlying effect size reduction was associated with outcome-relevant experimental differences between exploratory and confirmatory stages. In summary, multi-laboratory studies proved highly informative and served as an effective filter for promising treatments.

scientific communication and education↗

Technology-enhanced learning in undergraduate neuroscience education: tractography-based virtual dissection in psychology

Background: Neuroanatomy poses a significant challenge for Psychology students due to its spatial and conceptual complexity. Educational approaches that enhance the relevance and visualization of neuroanatomical content may improve students learning experiences. This study implemented a tractography-based activity focused on the virtual dissection of the arcuate fasciculus, a major white matter pathway, in undergraduate Psychology students and examined the relationships between perceived learning and students perceptions of utility, difficulty and handling, and organizational aspects of the activity. Methods: First-year undergraduate Psychology students participated in a two-session tractography-based activity combining instruction on white matter anatomy and diffusion tractography with a hands-on virtual dissection of the arcuate fasciculus using research-grade software routinely employed in neuroscience research. Following the activity, students completed an anonymous questionnaire assessing perceived learning, utility, difficulty and handling, and organizational aspects of the activity. Pearson correlations, multiple regression analyses, and relative importance analyses were performed. Results: Sixty-eight students completed the questionnaire. Students reported generally positive perceptions of the activity across the evaluated dimensions, with perceived learning receiving the highest mean score (M = 3.44, SD = .85). Perceived utility showed the strongest association with perceived learning (r = .64, p < .001). The regression model explained 41% of the variance in perceived learning (R2 = .41, adjusted R2 = .38, p < .001). Perceived utility was the only significant predictor in the model ({beta} = .59, p = .001), accounting for 67.9% of the explained variance. Conclusions: The findings support the feasibility of integrating authentic neuroimaging tools into undergraduate neuroanatomy teaching. Students who perceived the activity as more useful also reported higher perceived learning outcomes, with perceived utility emerging as the strongest predictor of perceived learning. In contrast, perceived difficulty and handling, and organizational aspects did not make significant independent contributions. These results suggest that students perceptions of educational relevance may play an important role in technology-enhanced STEM learning experiences.

scientific communication and education↗

A randomized trial of grant writing coaching groups: Qualitative interviews revealing key elements of intervention efficacy

Background Training in grant proposal writing is an essential component of professional development for academic scientists in the biomedical and behavioral sciences. Despite the expansion of inter- and intra-institutional grant writing coaching groups as an approach to honing these skills, specific features that enhance or limit coaching group effectiveness have not been rigorously studied. Methods Qualitative inerviews were conducted with a subset of early-career investigators (n=204 and coaches (n=36) engaged in a national, U.S. based, group-randomized trial of grant writing coaching groups to test the effects of two variables on submission and funding of national-level proposals: (1) coaching duration (regular/extended dose) and (2) mode of engaging a scientific advisor (someone with content-aligned expertise) in the coaching process. This report focuses on interviews conducted upon completion of the regular coaching dose - 5 months of biweekly, group-based coaching sessions to support active proposal writing. Interviews were designed to identify which coaching group elements were perceived to be the most critical. Transcribed interviews were analyzed using deductive (participants) or open (coaches) coding to identify themes. Results Triangulation of results from participant and coach interviews showed strong concordance that coaches, peers, and scientific advisors all played key roles in supporting intervention efficacy. Sufficient alignment of scientific fields and/or methodologies among group members was important, although breadth of perspectives was also valued. Other critical group features were skilled and well-organized coaches, detailed feedback, peer-to-peer support (technical and psychosocial), and clear expectations for group functioning. Factors attenuating impact included variation in participants' engagement and "readiness to write," within-group mismatches of expertise or grant mechanisms, and limited research support at some participants' home institutions. Conclusion This study identified key elements of successful grant writing coaching groups and potential barriers to their effectiveness, while yielding insights about tailoring this approach for individuals at different stages of proposal development.

scientific communication and education↗