bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Scientific Communication and Education”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Biomedical Text Mining for Research Rigor and Integrity: Tasks, Challenges, Directions

An estimated quarter of a trillion US dollars is invested in the biomedical research enterprise annually. There is growing alarm that a significant portion of this investment is wasted, due to problems in reproducibility of research findings and in the rigor and integrity of research conduct and reporting. Recent years have seen a flurry of activities focusing on standardization and guideline development to enhance the reproducibility and rigor of biomedical research. Research activity is primarily communicated via textual artifacts, ranging from grant applications to journal publications. These artifacts can be both the source and the end result of practices leading to research waste. For example, an article may describe a poorly designed experiment, or the authors may reach conclusions not supported by the evidence presented. In this article, we pose the question of whether biomedical text mining techniques can assist the stakeholders in the biomedical research enterprise in doing their part towards enhancing research integrity and rigor. In particular, we identify four key areas in which text mining techniques can make a significant contribution: plagiarism/fraud detection, ensuring adherence to reporting guidelines, managing information overload, and accurate citation/enhanced bibliometrics. We review the existing methods and tools for specific tasks, if they exist, or discuss relevant research that can provide guidance for future work. With the exponential increase in biomedical research output and the ability of text mining approaches to perform automatic tasks at large scale, we propose that such approaches can add checks and balances that promote responsible research practices and can provide significant benefits for the biomedical research enterprise.\n\nSupplementary informationSupplementary material is available at BioRxiv.

scientific communication and education

Unmet Needs for Analyzing Biological Big Data: A Survey of 704 NSF Principal Investigators

In a 2016 survey of 704 National Science Foundation (NSF) Biological Sciences Directorate principal investigators (BIO PIs), nearly 90% indicated they are currently or will soon be analyzing large data sets. BIO PIs considered a range of computational needs important to their work--including high performance computing (HPC), bioinformatics support, multi-step workflows, updated analysis software, and the ability to store, share, and publish data. Previous studies in the United States and Canada emphasized infrastructure needs. However, BIO PIs said the most pressing unmet needs are training in data integration, data management, and scaling analyses for HPC--acknowledging that data science skills will be required to build a deeper understanding of life. This portends a growing data knowledge gap in biology and challenges institutions and funding agencies to redouble their support for computational training in biology.

scientific communication and education

Reproducibility2020: Progress and Priorities

The preclinical research process is a cycle of idea generation, experimentation, and reporting of results. The biomedical research community relies on the reproducibility of published discoveries to create new lines of research and to translate research findings into therapeutic applications. Since 2012, when scientists from Amgen reported that they were able to reproduce only 6 of 53 \"landmark\" preclinical studies, the biomedical research community began discussing the scale of the reproducibility problem and developing initiatives to address critical challenges. GBSI released the \"Case for Standards\" in 2013, one of the first comprehensive reports to address the rising concern of irreproducible biomedical research. Further attention was drawn to issues that limit scientific self-correction including reporting and publication bias, underpowered studies, lack of open access to methods and data, and lack of clearly defined standards and guidelines in areas such as reagent validation. To evaluate the progress made towards reproducibility since 2013, GBSI identified and examined initiatives designed to advance quality and reproducibility. Through this process, we identified key roles for funders, journals, researchers and other stakeholders and recommended actions for future progress. This paper describes our findings and conclusions.

scientific communication and education

Maintaining the provenance of microscopy metadata using OMERO.forms software

The creation of datasets that are findable, accessible, interoperable and reproducible (the FAIR standard) requires that data provenance be maintained1. Provenance is particularly important for microscopy data, whose interpretation is dependent on the biological context (e.g. cell state) and detection reagent (e.g. antibody.) This paper describes a new software tool, OMERO.forms, that extends the OMERO microscopy data management system2 to simplify and enhance metadata entry and provenance tracking.

scientific communication and education

TOWARDS A GLOBAL SUPPORT OF CORE DATA RESOURCES FOR THE LIFE SCIENCES

On November 18-19, 2016, the Human Frontier Science Program Organization (HFSPO) hosted a meeting of senior managers of key data resources and leaders of several major funding organizations to discuss the challenges associated with sustaining biological and biomedical (i.e., life sciences) data resources and associated infrastructure. A strong consensus emerged from the group that core data resources for the life sciences should be supported through a coordinated international effort(s) that better ensure long-term sustainability and that appropriately align funding with scientific impact. Ideally, funding for such data resources should allow for access at no charge, as is presently the usual (and preferred) mechanism. Below, the rationale for this vision is described, and some important considerations for developing a new international funding model to support core data resources for the life sciences are presented.

scientific communication and education

Beyond the limitation of randomized controlled trials (RCTs)-current drug repositioning by using human induced pluripotent stem (iPS) cells technology-

To ensure the clinical value of medical interventions, randomized controlled trials (RCTs) are necessary. However, the results of conventional RCTs cannot show individual therapeutic efficacy and safety for medical intervention to a targeted patient. It is the most important weak point of conventional RCTs. Here I show that the new clinical research method by using human induced pluripotent stem (iPS) cells technology will be able to complement the most important weak point of conventional RCTs.\n\nAs the representative examples, I show the new clinical values of statins (inhibitors of 3-hydroxy-3-methylglutaryl-coenzyme A reductase) found by using human iPS cells technology in achondroplasia or hanatophoric dysplasia (type 1) case and hepatitis C virus (HCV) infection case. Furthermore, they are also important examples for drug repositioning.\n\nTherefore, my article would be valuable as a scientific communication for physicians and/or scientists.

scientific communication and education

Sustaining Scholarly Infrastructures through Collective Action: The lessons that Olson can teach us

Infrastructures for data, such as repositories, curation systems, aggregators, indexes and standards are public goods. This means that finding sustainable economic models to support them is a challenge. This is due to free-loading, where someone who does not contribute to the support of the infrastructure nonetheless gains the benefit of it. The work of Mancur Olson (1974) suggests there are only three ways to address this for large groups: compulsion (often as some form of taxation) to support the infrastructure; the provision of non-collective (club) goods limited to those who contribute as a side-effect of providing the collective good; or mechanisms that lower the effective number of participants in the negotiation (oligopoly).\n\nIn this paper I use Olsons framework to analyze existing scholarly infrastructures and proposals for the sustainability of new infrastructures. I argue that the focus on sustainability models prior to seeking a set of agreed governance principles is the wrong approach. Rather we need to understand how to navigate from club-like to public-like goods. We need to define the communities that contribute and identify club-like benefits for those contributors. We need interoperable principles of governance and resourcing to provide public-like goods and we should draw on the political economics of taxation to develop this.

scientific communication and education

Amending Published Articles: Time To Rethink Retractions And Corrections?

Academic publishing is evolving and our current system of correcting research post-publication is failing, both ideologically and practically. It does not encourage researchers to engage in consistent post-publication changes. Worse yet, post-publication updates are misconstrued as punishments or admissions of guilt. We propose a different model that publishers of research can apply to the content they publish, ensuring that any post-publication amendments are seamless, transparent and propagated to all the countless places online where descriptions of research appear. At the center, the neutral term \"amendment\" describes all forms of post-publication change to an article. We lay out a straightforward and consistent process that applies to each of the three types of amendments: insubstantial, substantial, and complete. This proposed system supports the dynamic nature of the research process itself as researchers continue to refine or extend the work, removing the emotive climate particularly associated with retractions and corrections to published work. It allows researchers to cite and share the correct versions of articles with certainty, and for decision makers to have access to the most up to date information.

scientific communication and education

The Readability Of Scientific Texts Is DecreasingOver Time

Clarity and accuracy of reporting are fundamental to the scientific process. The understandability of written language can be estimated using readability formulae. Here, in a corpus consisting of 707 452 scientific abstracts published between 1881 and 2015 from 122 influential biomedical journals, we show that the readability of science is steadily decreasing. Further, we demonstrate that this trend is indicative of a growing usage of general scientific jargon. These results are concerning for scientists and for the wider public, as they impact both the reproducibility and accessibility of research findings.

scientific communication and education

EXPLANe: An Extensible Framework for Poster Annotation with Mobile Devices

SummaryScientific posters tend to be brief, unstructured, and generally unsuitable for communication beyond a poster session. This paper describes EXPLANe, a framework for annotating posters using optical text recognition and web services on mobile devices. EXPLANe is demonstrated through an interface to the MyVariant.info variant annotation web services, and provides users a list of biological information linked with genetic variants (as found via extracted RSIDs from annotated posters). This paper delineates the architecture of the application, and includes results of a five-part evaluation we conducted. Researchers and developers can use the existing codebase as a foundation from which to generate their own annotation tabs when analyzing and annotating posters.\n\nAvailabilityAlpha EXPLANe software is available as an open source application at https://github.com/ngopal/EXPLANe\n\nContactSean D. Mooney (sdmooney@uw.edu)

scientific communication and education

Addressing the digital divide in contemporary biology: Lessons from teaching UNIX

Researchers in the biomedical sciences increasingly rely on applications that lack a graphical interface and require inputting code that, such as UNIX. Scientists who are not trained in computer science face an enormous challenge in analyzing the high-throughput data their research groups generate. We present a training model for use of command-line tools when the learner has little to no prior knowledge of UNIX.

scientific communication and education

Standardising and harmonising research data policy in scholarly publishing

Practice paper Practice paper References Research data policies influence researchers willingness to share research data to varying extents (Meadows, 2014; Schmidt, Gemeinholzer, & Treloar, 2016). A growing number of research funders and institutions are introducing policies on research data sharing. These include the National Institutes of Health (NIH), Gates Foundation, the EU Horizon 2020 programme, Wellcome Trust and the seven UK research councils (Hahnel, 2015). Policy requirements vary, with some requiring researchers to prepare data management plans and others, such as the Engineering and Physical Sciences Research Council (EPSRC), requiring evidence of public data archiving to be included in published research papers. To support publication of more reproducible research scholarly journals, societies and conferences ...

scientific communication and education

Looking into Pandora’s Box: The Content of Sci-Hub and its Usage

Despite the growth of Open Access, illegally circumventing paywalls to access scholarly publications is becoming a more mainstream phenomenon. The web service Sci-Hub is amongst the biggest facilitators of this, offering free access to around 62 million publications. So far it is not well studied how and why its users are accessing publications through Sci-Hub. By utilizing the recently released corpus of Sci-Hub and comparing it to the data of {small tilde}28 million downloads done through the service, this study tries to address some of these questions. The comparative analysis shows that both the usage and complete corpus is largely made up of recently published articles, with users disproportionately favoring newer articles and 35% of downloaded articles being published after 2013. These results hint that embargo periods before publications become Open Access are frequently circumnavigated using Guerilla Open Access approaches like Sci-Hub. On a journal level, the downloads show a bias towards some scholarly disciplines, especially Chemistry, suggesting increased barriers to access for these. Comparing the use and corpus on a publisher level, it becomes clear that only 11% of publishers are highly requested in comparison to the baseline frequency, while 45% of all publishers are significantly less accessed than expected. Despite this, the oligopoly of publishers is even more remarkable on the level of content consumption, with 80% of all downloads being published through only 9 publishers. All of this suggests that Sci-Hub is used by different populations and for a number of different reasons, and that there is still a lack of access to the published scientific record. A further analysis of these openly available data resources will undoubtedly be valuable for the investigation of academic publishing.

scientific communication and education

Why Do Scientists Fabricate And Falsify Data? A Matched-Control Analysis Of Papers Containing Problematic Image Duplications

It is commonly hypothesized that scientists are more likely to engage in data falsification and fabrication when they are subject to pressures to publish, when they are not restrained by forms of social control, when they work in countries lacking policies to tackle scientific misconduct, and when they are male. Evidence to test these hypotheses, however, is inconclusive due to the difficulties of obtaining unbiased data.\n\nHere we report a pre-registered test of these four hypotheses, conducted on papers that were identified in a previous study as containing problematic image duplications through a systematic screening of the journal PLoS ONE. Image duplications were classified into three categories based on their complexity, with category 1 being most likely to reflect unintentional error and category 3 being most likely to reflect intentional fabrication. Multiple parameters connected to the hypotheses above were tested with a matched-control paradigm, by collecting two controls for each paper containing duplications.\n\nCategory 1 duplications were mostly not associated with any of the parameters tested, in accordance with the assumption that these duplications were mostly not due to misconduct. Category 2 and 3, however, exhibited numerous statistically significant associations. Results of univariable and multivariable analyses support the hypotheses that academic culture, peer control, cash-based publication incentives and national misconduct policies might affect scientific integrity. Significant correlations between the risk of image duplication and individual publication rates or gender, however, were only observed in secondary and exploratory analyses.\n\nCountry-level parameters generally exhibited effects of larger magnitude than individual-level parameters, because a subset of countries was significantly more likely to produce problematic image duplications. Promoting good research practices in all countries should be a priority for the international research integrity agenda.

scientific communication and education

Anticipated effects of an open access policy at a private foundation

BackgroundThe Gordon and Betty Moore Foundation (GBMF) was interested in understanding the potential effects of a policy requiring open access to peer-reviewed publications resulting from the research the foundation funds.\n\nMethodsWe collected data on more than 2000 publications in over 500 journals that were generated by GBMF grantees since 2001. We then examined the journal policies to establish how two possible open access policies might have affected grantee publishing habits.\n\nResultsWe found that 99.3% of the articles published by grantees would have complied with a policy that requires open access within 12 months of publication. We also estimated the maximum annual costs to GBMF for covering fees associated with \"gold open access\" to be between $400,000 and $2,600,000 annually.\n\nDiscussionBased in part on this study, GBMF has implemented a new open access policy that requires grantees make peer-reviewed publications fully available within 12 months.

scientific communication and education

A persistent lack of International representation on editorial boards in biology

The scholars comprising journal editorial boards play a critical role in defining the trajectory of knowledge in their field. Nevertheless, studies of editorial board composition remain rare, especially those focusing on journals publishing research in the increasingly globalized fields of science, technology, engineering, and math (STEM). Using metrics for quantifying the diversity of ecological communities, we quantified international representation on the 1985-2014 editorial boards of twenty-four environmental biology journals. Over the course of three decades there were 3831 unique scientists based in 70 countries that served as editors. The size of the editorial community increased over time - there were 420% more editors serving in 2014 than in 1985 - as did the number of countries in which editors were based. Nevertheless, editors based outside the Global North (the group of economically developed countries with high per capita Gross Domestic Product (GDP) that collectively concentrate most global wealth) were extremely rare. Furthermore, 67.06% of all editors were based in either the USA or UK. Consequently, Geographic Diversity - already low in 1985 - remained unchanged through 2014. We argue that this limited geographic diversity can detrimentally affect the creativity of scholarship published in journals, the progress and direction of research, the composition of the STEM workforce, and the development of science in Latin America, Africa, the Middle East, and much of Asia (i.e., the Global South).

scientific communication and education

A Universal Metric For Evaluating, Optimising And Benchmarking The Performance Of A Research Technology Platform (RTP)

Research Technology Platforms (RTPs) exist to facilitate the application and utilisation of specific analytical technologies to the highest possible standard thus delivering reputable data across a broad spectrum of research themes. Specifically, RTPs centralise expertise in a given technology and provide an unparalleled level of continuity and practical knowledge retention that simply cannot be achieved by more organic, ad hoc means of support. As small non profit businesses often tasked with recovering all or a percentage of their running costs, RTPs are under significant pressure to keep pace with rapidly advancing technology and new methodologies against a back drop of dwindling funding for scientific research. At present there are a number of non-trivial issues that make assessing the operational performance of a RTP difficult to determine on a standalone basis let alone attempting to benchmark against other RTPs within the same or different technology fields. Firstly, depending on the technological speciality the RTP may work to one of essentially three operational models. RTPs such as Bio-Imaging or Cytometry provide access to well-maintained analytical systems that can be utilised by trained individuals for a timed access charge. In some cases there will be a requirement for assisted operation of certain instruments by core staff (e.g. cell sorters). Genomics and Proteomics RTPs tend to function on a project basis whereby users will not access the technology themselves rather pay for a full analytical service often with a milestone-based approach for tracking progress. Other RTPs work to a hybrid approach were technical staff provide certain elements of sample preparation for specific projects prior to analysis on core supported, user accessible instrumentation. Secondly the specific operational costs that each RTP is tasked to recover varies significantly on a local, national and international level due to institutional subsidies. These operational costs can include staff salaries, instrument maintenance, associated running consumables, and in some cases instrument depreciation but there is standardised rule as to what each RTP is tasked to recover and to what percentage.\n\nHere we present a generalised mathematical approach to describe the customisable metrics of any given RTP serviceThe general strategy how to increase performance within the framework of this approach has been identified through breaking down these customisable metrics into components and maximising them according to specific requirements. These strategies could be potentially adopted for different operational or local procedures, integrating the specifics related to the institutional or national policies. The approach laid down here should be considered as a trigger for opening a discussion around how to address optimising RTP performance and allow for benchmarking across the full breadth of RTPs.

scientific communication and education

Research Waste In ME/CFS

ObjectiveTo compare the prevalence of selective reporting in ME/CFS research areas: psychosocial versus cellular.\n\nMethodA bias appraisal was conducted on three trials (1x psychosocial and 2x cellular) to compare risk of bias in study design, selection and measurement. The primary outcome compared evidence and justifications in resolving biases by proportions (%) and ORs (Odds Ratio); the secondary outcome determined the proportion (in %) of ME/CFS grants at risk of bias.\n\nResultsNS (cellular study) was twice as likely to present evidence in resolving biases over PACE (psychosocial trial) (OR = 2.16; 65.6% vs 46. 9%), but this difference was not significant (p = 0.13). However, NS was five times more likely to justify biases over PACE (OR = 4.76; 46.9% vs 15. 6%) and this difference was significant (p = 0.0095; p < 0.05). PACE was weak in place (operational aspects 32%) and NS for data practices (37%). The proportion of grants were more biased in PACE (72%) than NS (28%) for evidence, and also more biased in PACE (86%) than NS (14%) for justifications.\n\nConclusionPsychosocial trials on ME/CFS are more likely to engage in selective reporting indicative of research waste than cellular trials. Improvements to place may help reduce these biases, whereas cellular trials may benefit from adopting more translatable data methods. However, these findings are based on two trials. Further risk of bias appraisals are needed to determine the number of trials required to make robust these findings.

scientific communication and education