bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.03.27.011106

Scientometric correlates of high-quality reference lists in ecological papers

Abstract

It is said that the quality of a scientific publication is as good as the science it cites, but the properties of high-quality reference lists have never been numerically quantified. We examined seven numerical characteristics of reference lists of 50,878 primary research articles published in 17 ecological journals between 1997 and 2017. Over this 20-years period, there have been significant changes in reference lists’ properties. On average, more recent ecological papers have longer reference lists, cite more high Impact Factor papers, and fewer non-journal publications. Furthermore, we show that highly cited papers across the ecology literature have longer reference lists, cite more recent and impactful papers, and account for more self-citations. Conversely, the proportion of ‘classic’ papers and non-journal publications cited, as well as the temporal range of the reference list, have no significant influence on articles’ citations. From this analysis, we distill a recipe for crafting impactful reference lists.View Full Text

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mammola, S., Martinez, A., Fontaneto, D., Chichorro, F.. 2020-03-29. Scientometric correlates of high-quality reference lists in ecological papers. https://doi.org/10.1101/2020.03.27.011106

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Dynamics of the COVID -19 Related Publications

BackgroundThis study aims to analyze the dynamics of the published articles and preprints of Covid-19 related literature from different scientific databases and sharing platforms. MethodsThe PubMed, Elsevier, and Research Gate (RG) databases were under consideration in this study over a specific time. Analyses were carried out on the number of publications as (a) function of time (day), (b) journals and (c) authors. Doubling time of the number of publications was analyzed for PubMed "all articles" and Elsevier published articles. Analyzed databases were (1A) PubMed "all articles" (01/12/2019-12/06/2020) (1B) PubMed Review articles (01/12/2019-2/5/2020) and (1C) PubMed Clinical Trials (01/01/2020-30/06/2020) (2) Elsevier all publications (01/12/2019-25/05/2020) (3) RG (Article, Pre Print, Technical Report) (15/04/2020-30/4/2020). FindingsTotal publications in the observation period for PubMed, Elsevier, and RG were 23000, 5898 and 5393 respectively. The average number of publications/day for PubMed, Elsevier and RG were 70.0 {+/-}128.6, 77.6{+/-}125.3 and 255.6{+/-}205.8 respectively. PubMed shows an avalanche in the number of publication around May 10, number of publications jumped from 6.0{+/-}8.4/day to 282.5{+/-}110.3/day. The average doubling time for PubMed, Elsevier, and RG was 10.3{+/-}4 days, 20.6 days, and 2.3{+/-}2.0 days respectively. In PubMed average articles/journal was 5.2{+/-}10.3 and top 20 authors representing 935 articles are of Chinese descent. The average number of publications per author for PubMed, Elsevier, and RG was 1.2{+/-}1.4, 1.3{+/-}0.9, and 1.1{+/-}0.4 respectively. Subgroup analysis, PubMed review articles mean and median review time for each article were <0|17{+/-}17|77> and 13.9 days respectively; and reducing at a rate of-0.21 days (count)/day. InterpretationAlthough the disease has been known for around 6 months, the number of publications related to the Covid-19 until now is huge and growing very fast with time. It is essential to rationalize the publications scientifically by the researchers, authors, reviewers, and publishing houses. FundingNone

scientific communication and education

From bioinformatics user to bioinformatics engineer: a report

Teaching computer programming is not a simple task and it is challenging to introduce the concepts of programming in graduate programs of other fields. Little efforts have been made on engaging students in computational development after programming trainings. An emerging need is to establish subjects of bioinformatics and programming languages in genetics and molecular biology graduate programs, when students in these degree programs are immersed in a sea of genomic and transcriptomic data, which demands proficient computational treatment. I report an empirical guideline to introduce programming languages and recommend Python as first language for graduate programs in which students were from genetics and molecular biology backgrounds. Including the development of programming solutions related to graduate students' research activities may improve programming skills and better engagement. These results suggest that the applied approach leads to enhanced learning of introductory to autonomy in highly advanced programming concepts by graduate students. This guide should be extended for other research programs.

scientific communication and education