bioRxiv · 10.1101/247775
Reproducible, flexible and high throughput data extraction from primary literature: The metaDigitise R package
Abstract
O_LIResearch synthesis, such as comparative and meta-analyses, requires the extraction of effect sizes from primary literature, which are commonly calculated from descriptive statistics. However, the exact values of such statistics are commonly hidden in figures.\nC_LIO_LIExtracting descriptive statistics from figures can be a slow process that is not easily reproducible. Additionally, current software lacks an ability to incorporate important meta-data (e.g., sample sizes, treatment / variable names) about experiments and is not integrated with other software to streamline analysis pipelines.\nC_LIO_LIHere we present the R package metaDigitise which extracts descriptive statistics such as means, standard deviations and correlations from four plot types: 1) mean/error plots (e.g. bar graphs with standard errors), 2) box plots, 3) scatter plots and 4) histograms. metaDigitise is user-friendly and easy to learn as it interactively guides the user through the data extraction process. Notably, it enables large-scale extraction by automatically loading image files, letting the user stop processing, edit and add to the resulting data-frame at any point.\nC_LIO_LIDigitised data can be easily re-plotted and checked, facilitating reproducible data extraction from plots with little inter-observer bias. We hope that by making the process of figure extraction more flexible and easy to conduct it will improve the transparency and quality of meta-analyses in the future.\nC_LI
Source connections
Explore related subjects
Keep this discovery
Pick, J. L., Nakagawa, S. W., Noble, D. W.. 2018-01-15. Reproducible, flexible and high throughput data extraction from primary literature: The metaDigitise R package. https://doi.org/10.1101/247775
Cite the original work for its findings. Save a collection to share your selection of sources.