bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.09.23.310052

Co-Citation Percentile Rank and JYUcite: a new network-standardized output-level citation influence metric and its implementation using Dimensions API

Abstract

Judging value of scholarly outputs quantitatively remains a difficult but unavoidable challenge. Most of the proposed solutions suffer from three fundamental shortcomings: they involve i) the concept of journal, in one way or another, ii) calculating arithmetic averages from extremely skewed distributions, and iii) binning data by calendar year. Here, we introduce a new metric Co-citation Percentile Rank (CPR), that relates the current citation rate of the target output taken at resolution of days since first citable, to the distribution of current citation rates of outputs in its co-citation set, as its percentile rank in that set. We explore some of its properties with an example dataset of all scholarly outputs from University of Jyvaskyla spanning multiple years and disciplines. We also demonstrate how CPR can be efficiently implemented with Dimensions database API, and provide a publicly available web resource JYUcite, allowing anyone to retrieve CPR value for any output that has a DOI and is indexed in the Dimensions database. Finally, we discuss how CPR remedies failures of the Relative Citation Ratio (RCR), and remaining issues in situations where CPR too could potentially lead to biased judgement of value.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Seppanen, J.-T., Varri, H., Ylonen, I.. 2020-09-23. Co-Citation Percentile Rank and JYUcite: a new network-standardized output-level citation influence metric and its implementation using Dimensions API. https://doi.org/10.1101/2020.09.23.310052

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Applying And Promoting Open Science In Ecology - Surveyed Drivers And Challenges

Open Science (OS) comprises a variety of practices and principles that are broadly intended to improve the quality and transparency of research, and the concept is gaining traction. Since OS has multiple facets and still lacks a unifying definition, it may be interpreted quite differently among practitioners. Moreover, successfully implementing OS broadly throughout science requires a better understanding of the conditions that facilitate or hinder OS engagement, and in particular, how practitioners learn OS in the first place. We addressed these issues by surveying OS practitioners that attended a workshop hosted by the Living Norway Ecological Data Network in 2020. The survey contained scaled-response and open-ended questions, allowing for a mixed-methods approach. Out of 128 registered participants we obtained survey responses from 60 individuals. Responses indicated usage and sharing of data and code, as well as open access publications, as the OS aspects most frequently engaged with. Men and those affiliated with academic institutions reported more frequent engagement with OS than women and those with other affiliations. When it came to learning OS practices, only a minority of respondents reported having encountered OS in their own formal education. Consistent with this, a majority of respondents viewed OS as less important in their teaching than in their research and supervision. Even so, many of the respondents suggestions for what would help or hinder individual OS engagement included more knowledge, guidelines, resource availability and social and structural support; indicating that formal instruction can facilitate individual OS engagement. We suggest that the time is ripe to incorporate OS in teaching and learning, as this can yield substantial benefits to OS practitioners, student learning, and ultimately, the objectives advanced by the OS movement.

scientific communication and education

An exploratory analysis of 4844 withdrawn articles and their retraction notes.

The objective of our study was to obtain an updated image of the dynamic of retractions and retraction notes, retraction reasons for questionable research and publication practices, countries producing retracted articles, and the scientific impact of retractions by studying 4844 PubMed indexed retracted articles published between 2009 and 2020 and their retraction notes. RESULTSMistakes/inconsistent data account for 32% of total retractions, followed by images(22,5%), plagiarism(13,7%) and overlap(11,5%). Thirty countries account for 94,79% of 4844 retractions. Top five are: China(32,78%), United States(18,84%), India(7,25%), Japan(4,37%) and Italy(3,75%). The total citations number for all articles is 140810(Google Scholar), 96000(Dimensions). Average exposure time(ET) is 28,89 months. Largest ET is for image retractions(49,3 months), lowest ET is for editorial errors(11,2 months). The impact of retracted research is higher for Spain, Sweden, United Kingdom, United States, and other nine countries and lower for Pakistan, Turkey, Malaysia, and other six countries, including China. CONCLUSIONSMistakes and data inconsistencies represent the main retraction reason; images and ethical issues show a growing trend, while plagiarism and overlap still represent a significant problem. There is a steady increase in QRP and QPP article withdrawals. Retraction of articles seems to be a technology-dependent process. The number of citations of retracted articles shows a high impact of papers published by authors from certain countries. The number of retracted articles per country does not always accurately reflect the scientific impact of QRP/QPP articles. The country distribution of retraction reasons shows structural problems in the organization and quality control of scientific research, which have different images depending on geographical location, economic development, and cultural model.

scientific communication and education