bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.11.07.565991

ABCal: a Python package for Author Bias Computation and Scientometric Plotting for Reviews and Meta-Analyses

Abstract

Systematic reviews are critical summaries of the exiting literature on a given subject and, when combined with meta-analysis, provides a quantitative synthesis of evidence to direct and inform future research. Such reviews must, however, account for complex sources of between study heterogeneity and possible sources of bias, such as publication bias. This paper presents the methods and results of a research study using a newly developed software tool called ABCal (version 1.0.2) to compute and assess author bias in the literature, providing a quantitative measure for the possible effect of overrepresented authors introducing bias to the overall interpretation of the literature. ABCal includes a new metric referred to as author bias, which is a measure of potential biases per paper when the frequency or proportions of contributions from specific authors are considered. The metric is able to account for a significant portion of the observed heterogeneity between studies included in meta-analyses. A meta-regression between observed effect measures and author bias values revealed that higher levels of author bias were associated with higher effect measures while lower author bias was evident for studies with lower effect measures. Furthermore, the softwares capabilities to analyse authorship contributions and produce scientometric plots was able to reveal distinct patterns in both the temporal and geographic distributions of publications, which may relate to any evident publication bias. Thus, ABCal can aid researchers in gaining a deeper understanding of the research landscape and assist in identifying both key contributors and holistic research trends.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Le Clercq, L. S.. 2023-11-09. ABCal: a Python package for Author Bias Computation and Scientometric Plotting for Reviews and Meta-Analyses. https://doi.org/10.1101/2023.11.07.565991

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

A Large Language Model-Powered Map of Metabolomics Research

We present a comprehensive map of the metabolomics research landscape, synthesizing insights from over 80,000 publications. Using PubMedBERT, we transformed abstracts into 768-dimensional embeddings that capture the nuanced thematic structure of the field. Dimensionality reduction with t-SNE revealed distinct clusters corresponding to key domains such as analytical chemistry, plant biology, pharmacology, and clinical diagnostics. In addition, a neural topic modeling pipeline refined with GPT-4o mini reclassified the corpus into 20 distinct topics--ranging from "Plant Stress Response Mechanisms" and "NMR Spectroscopy Innovations" to "COVID-19 Metabolomic and Immune Responses." Temporal analyses further highlight trends including the rise of deep learning methods post-2015 and a continued focus on biomarker discovery. Integration of metadata such as publication statistics and sample sizes provide additional context to these evolving research dynamics. An interactive web application (https://metascape.streamlit.app/) enables dynamic exploration of these insights. Overall, this study offers a robust framework for literature synthesis that empowers researchers, clinicians, and policymakers to identify emerging research trajectories and address critical challenges in metabolomics, while also sharing our perspectives on key trends shaping the field.

scientific communication and education↗

Addressing cultural and knowledge barriers to enable preclinical sex inclusive research

For over thirty years, research has highlighted a sex bias in early research, risking the validity of biological knowledge. The first step towards change is effectively challenging misconceptions allowing researchers to perceive sex inclusive research as do-able. Utilising the theory of planned behaviour, we quantified researchers intention as a proxy measure for conducting sex inclusive research and explored attitude (value of the behaviour), subjective norm (perceived social pressure) and behavioural control (ability to conduct the behaviour). Additionally, we quantified the knowledge gap, prevalence of misconceptions, and assessed perceived benefits and barriers. We tested a workshop intervention that directly challenges the cultural embedded barriers. The data shows researchers intentions were high, but they had weak statistical knowledge and misunderstandings leading to a perception that inclusive research is prohibitive due to cost and animal use. We demonstrate that participation in the training intervention improved knowledge, altered the perceived barriers and cultural expectations.

scientific communication and education↗