bioRxiv ScienceSearch

Biology subjects

Macleod, M. R.

Publications and source records attributed to Macleod, M. R..

7 recordsLinked to original sources

A randomised controlled trial of an Intervention to Improve Compliance with the ARRIVE guidelines (IICARus)

The ARRIVE (Animal Research: Reporting of In Vivo Experiments) guidelines are widely endorsed but compliance is limited. We sought to determine whether journal-requested completion of an ARRIVE checklist improves full compliance with the guidelines. In a randomised controlled trial, manuscripts reporting in vivo animal research submitted to PLOS ONE (March-June 2015) were allocated to either requested completion of an ARRIVE checklist or current standard practice. We measured the change in proportion of manuscripts meeting all ARRIVE guideline checklist items between groups. We randomised 1,689 manuscripts, 1,269 were sent for peer review and 762 accepted for publication. The request to complete an ARRIVE checklist had no effect on full compliance with the ARRIVE guidelines. Details of animal husbandry (ARRIVE sub-item 9a) was the only item to show improved reporting, from 52.1% to 74.1% (X2=34.0, df=1, p=2.1x10-7). These results suggest that other approaches are required to secure greater implementation of the ARRIVE guidelines.\n\nBackgroundThere are widespread failures across in vivo animal research to adequately describe and report research methods, including critical measures to reduce the risk of experimental bias (Kilkenny et al., 2009, Macleod et al., 2015). Such omissions have been shown to be associated with overestimation of effect sizes (Macleod et al., 2015, Hirst et al., 2014) and are likely to contribute, in part, to translational failure. In an effort to improve reporting standards, an expert working group coordinated by the National Centre for the Replacement, Refinement and Reduction of Animals in Research (NC3Rs) developed the Animal Research: Reporting of In Vivo Experiments (ARRIVE) guidelines (Kilkenny et al., 2010), published in 2010.\n\nSince the ARRIVE guidelines were first published, they have been endorsed by many journals in their instructions to authors, but this has not been accompanied by substantial improvements in reporting (Baker et al., 2014, McGrath and Lilley, 2015, Gulin et al., 2015a, Avey et al., 2016). Simply endorsing the guidelines does not appear to be sufficient to encourage compliance. Recent findings suggest that following the introduction of mandated completion of a distinct reporting checklist at ten Nature Journals at the stage of first revision significantly improved the quality in reporting versus that of comparator journals (Han et al., 2017, Macleod and The NPQIP Collaborative Group, 2017)\n\nPLOS ONE is an open access online only journal which at the time this study began published around 32,000 research articles per year. Of these, some 5,000 described in vivo research. At present, PLOS ONE instructions to authors encourage compliance with the ARRIVE guidelines, but do not mandate checklist completion. Journals have an important role to play in ensuring that the quality of reporting in the research they publish is robust, yet the most effective mechanism by which they can achieve this remains unclear.\n\nOur aim was to test the impact on the quality of published reports of an intervention which would request, at the time of manuscript submission, that authors complete a checklist detailing where in the manuscript the various components of the ARRIVE checklist were met. This study, to our knowledge, is the first randomised controlled trial of requested ARRIVE guideline completion.

scientific communication and education

Animal models of chemotherapy-induced peripheral neuropathy: a machine-assisted systematic review and meta-analysis A comprehensive summary of the field to inform robust experimental design

Background and aimsChemotherapy-induced peripheral neuropathy (CIPN) can be a severely disabling side-effect of commonly used cancer chemotherapeutics, requiring cessation or dose reduction, impacting on survival and quality of life. Our aim was to conduct a systematic review and meta-analysis of research using animal models of CIPN to inform robust experimental design.\n\nMethodsWe systematically searched 5 online databases (PubMed, Web of Science, Citation Index, Biosis Previews and Embase (September 2012) to identify publications reporting in vivo CIPN modelling. Due to the number of publications and high accrual rate of new studies, we ran an updated search November 2015, using machine-learning and text mining to identify relevant studies.\n\nAll data were abstracted by two independent reviewers. For each comparison we calculated a standardised mean difference effect size then combined effects in a random effects meta- analysis. The impact of study design factors and reporting of measures to reduce the risk of bias was assessed. We ran power analysis for the most commonly reported behavioural tests.\n\nResults341 publications were included. The majority (84%) of studies reported using male animals to model CIPN; the most commonly reported strain was Sprague Dawley rat. In modelling experiments, Vincristine was associated with the greatest increase in pain-related behaviour (-3.22 SD [-3.88; -2.56], n=152, p=0). The most commonly reported outcome measure was evoked limb withdrawal to mechanical monofilaments. Pain-related complex behaviours were rarely reported. The number of animals required to obtain 80% power with a significance level of 0.05 varied substantially across behavioural tests. Overall, studies were at moderate risk of bias, with modest reporting of measures to reduce the risk of bias.\n\nConclusionsHere we provide a comprehensive summary of the field of animal models of CIPN and inform robust experimental design by highlighting measures to increase the internal and external validity of studies using animal models of CIPN. Power calculations and other factors, such as clinical relevance, should inform the choice of outcome measure in study design.

neuroscience

Automation of citation screening in pre-clinical systematic reviews

BackgroundThe amount of published in vivo studies and the speed researchers are publishing them make it virtually impossible to follow the recent development in the field. Systematic review emerged as a method to summarise and analyse the studies quantitatively and critically but it is often out-of-date due to its lengthy process. MethodWe invited five machine learning and text-mining groups to build classifiers for identifying publications relevant to neuropathic pain (33814 training publications). We kept 1188 publications for the assessment of the performance of different classifiers. Two groups participated in the next stage: testing their algorithm on datasets labeled for psychosis (11777/2944) and datasets labeled for Vitamin D in multiple sclerosis (train/text: 2038/510). ResultThe performances (sensitive/specificity) of the most promising classifier built for neuropathic pain are: 95%/84%. The performance for psychosis and Vitamin D in multiple sclerosis datasets are 95%/73% and 100%/45%. ConclusionsMachine learning can significantly reduce the irrelevant publications in a systematic review, and save the scientists time and money. Classifier algorithms built for one dataset can be reapplied on another dataset in different field. We are building a machine learning service at the back of Systematic Review & Meta-analysis Facility (SyRF).

scientific communication and education

The use of text-mining and machine learning algorithms in systematic reviews: reducing workload in preclinical biomedical sciences and reducing human screening error

BackgroundHere we outline a method of applying existing machine learning (ML) approaches to aid citation screening in an on-going broad and shallow systematic review of preclinical animal studies, with the aim of achieving a high performing algorithm comparable to human screening.\n\nMethodsWe applied ML approaches to a broad systematic review of animal models of depression at the citation screening stage. We tested two independently developed ML approaches which used different classification models and feature sets. We recorded the performance of the ML approaches on an unseen validation set of papers using sensitivity, specificity and accuracy. We aimed to achieve 95% sensitivity and to maximise specificity. The classification model providing the most accurate predictions was applied to the remaining unseen records in the dataset and will be used in the next stage of the preclinical biomedical sciences systematic review. We used a cross validation technique to assign ML inclusion likelihood scores to the human screened records, to identify potential errors made during the human screening process (error analysis).\n\nResultsML approaches reached 98.7% sensitivity based on learning from a training set of 5749 records, with an inclusion prevalence of 13.2%. The highest level of specificity reached was 86%. Performance was assessed on an independent validation dataset. Human errors in the training and validation sets were successfully identified using assigned the inclusion likelihood from the ML model to highlight discrepancies. Training the ML algorithm on the corrected dataset improved the specificity of the algorithm without compromising sensitivity. Error analysis correction leads to a 3% improvement in sensitivity and specificity, which increases precision and accuracy of the ML algorithm.\n\nConclusionsThis work has confirmed the performance and application of ML algorithms for screening in systematic reviews of preclinical animal studies. It has highlighted the novel use of ML algorithms to identify human error. This needs to be confirmed in other reviews, , but represents a promising approach to integrating human decisions and automation in systematic review methodology.

neuroscience

Estimating the statistical performance of different approaches to meta-analysis of data from animal studies in identifying the impact of aspects of study design

BackgroundMeta-analysis is increasingly used to summarise the findings identified in systematic reviews of animal studies modelling human disease. Such reviews typically identify a large number of individually small studies, testing efficacy under a variety of conditions. This leads to substantial heterogeneity, and identifying potential sources of this heterogeneity is an important function of such analyses. However, the statistical performance of different approaches (normalised compared with standardised mean difference estimates of effect size; stratified meta-analysis compared with meta-regression) is not known.\n\nMethodsUsing data from 3116 experiments in focal cerebral ischaemia to construct a linear model predicting observed improvement in outcome contingent on 25 independent variables. We used stochastic simulation to attribute these variables to simulated studies according to their prevalence. To ascertain the ability to detect an effect of a given variable we introduced in addition this \"variable of interest\" of given prevalence and effect. To establish any impact of a latent variable on the apparent influence of the variable of interest we also introduced a \"latent confounding variable\" with given prevalence and effect, and allowed the prevalence of the variable of interest to be different in the presence and absence of the latent variable.\n\nResultsGenerally, the normalised mean difference (NMD) approach had higher statistical power than the standardised mean difference (SMD) approach. Even when the effect size and the number of studies contributing to the meta-analysis was small, there was good statistical power to detect the overall effect, with a low false positive rate. For detecting an effect of the variable of interest, stratified meta-analysis was associated with a substantial false positive rate with NMD estimates of effect size, while using an SMD estimate of effect size had very low statistical power. Univariate and multivariable meta-regression performed substantially better, with low false positive rate for both NMD and SMD approaches; power was higher for NMD than for SMD. The presence or absence of a latent confounding variables only introduced an apparent effect of the variable of interest when there was substantial asymmetry in the prevalence of the variable of interest in the presence or absence of the confounding variable.\n\nConclusionsIn meta-analysis of data from animal studies, NMD estimates of effect size should be used in preference to SMD estimates, and meta-regression should, where possible, be chosen over stratified meta-analysis. The power to detect the influence of the variable of interest depends on the effect of the variable of interest and its prevalence, but unless effects are very large adequate power is only achieved once at least 100 experiments are included in the meta-analysis.

bioinformatics

Findings of a retrospective, controlled cohort study of the impact of a change in Nature journals' editorial policy for life sciences research on the completeness of reporting study design and execution

ObjectiveTo determine whether a change in editorial policy, including the implementation of a checklist, has been associated with improved reporting of measures which might reduce the risk of bias.\n\nMethodsThe study protocol has been published at DOI: 10.1007/s11192-016-1964-8.\n\nDesignObservational cohort study.\n\nPopulationArticles describing research in the life sciences published in Nature journals, submitted after May 1st 2013.\n\nInterventionMandatory completion of a checklist at the point of manuscript revision.\n\nComparators(1) Articles describing research in the life sciences published in Nature journals, submitted before May 2013; (2) Similar articles in other journals matched for date and topic.\n\nPrimary OutcomeChange in proportion of Nature publications describing in vivo research published before and after May 2013 reporting the Landis 4 items (randomisation, blinding, sample size calculation, exclusions).\n\nWe included 448 NPG papers (223 published before May 2013, 225 after) identified by an individual hired by NPG for this specific task, working to a standard procedure; and an independent investigator used Pubmed Related Citations to identify 448 non-NPG papers with a similar topic and date of publication in other journals; and then redacted all publications for time sensitive information and journal name. Redacted manuscripts were assessed by 2 trained reviewers against a 74 item checklist, with discrepancies resolved by a third.\n\nResults394 NPG and 353 matching non-NPG publications described in vivo research. The number of NPG publications meeting all relevant Landis 4 criteria increased from 0/203 prior to May 2013 to 31/181 (16.4%) after (2-sample test for equality of proportions without continuity correction, X2 = 36.2, df = 1, p = 1.8 x 10-9). There was no change in the proportion of non- NPG publications meeting all relevant Landis 4 criteria (1/164 before, 1/189 after). There were more substantial improvements in the individual prevalences of reporting of randomisation, blinding, exclusions and sample size calculations for in vivo experiments, and less substantial improvements for in vitro experiments.\n\nConclusionsThere was a substantial improvement in the reporting of risks of bias in in vivo research in NPG journals following a change in editorial policy, to a level that to our knowledge has not been previously observed. However, there remain opportunities for further improvement.

scientific communication and education

Effect size and statistical power in the rodent fear conditioning literature - a systematic review

Proposals to increase research reproducibility frequently call for focusing on effect sizes instead of p values, as well as for increasing the statistical power of experiments. However, it is unclear to what extent these two concepts are indeed taken into account in basic biomedical science. To study this in a real-case scenario, we performed a systematic review of effect sizes and statistical power in studies on learning of rodent fear conditioning, a widely used behavioral task to evaluate memory. Our search criteria yielded 410 experiments comparing control and treated groups in 122 articles. Interventions had a mean effect size of 29.5%, and amnesia caused by memory-impairing interventions was nearly always partial. Mean statistical power to detect the average effect size observed in well-powered experiments with significant differences (37.2%) was 65%, and was lower among studies with non-significant results. Only one article reported a sample size calculation, and our estimated sample size to achieve 80% power considering typical effect sizes and variances (15 animals per group) was reached in only 12.2% of experiments. Actual effect sizes correlated with effect size inferences made by readers on the basis of textual descriptions of results only when findings were non-significant, and neither effect size nor power correlated with study quality indicators, number of citations or impact factor of the publishing journal. In summary, effect sizes and statistical power have a wide distribution in the rodent fear conditioning literature, but do not seem to have a large influence on how results are described or cited. Failure to take these concepts into consideration might limit attempts to improve reproducibility in this field of science.

neuroscience