bioRxiv ScienceSearch

bioRxiv · 10.1101/2020.04.11.037093

BIP4COVID19: Releasing impact metrics data for articles relevant to COVID-19

Abstract

Since the beginning of the 2019-20 coronavirus pandemic, a large number of relevant articles has been published or become available in preprint servers. These articles, along with earlier related literature, compose a valuable knowledge base affecting contemporary research studies, or even government actions to limit the spread of the disease and treatment decisions taken by physicians. However, the number of such articles is increasing at an intense rate making the exploration of the relevant literature and the identification of useful knowledge in it challenging. In this work, we describe BIP4COVID19, an open dataset compiled to facilitate the coronavirus-related literature exploration, by providing various indicators of scientific impact for the relevant articles. Additionally, we provide a publicly accessible Web interface on top of our data, allowing the exploration of the publications based on the computed indicators.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vergoulis, T., Kanellos, I., Chatzopoulos, S., Pla Karidi, D., Dalamagas, T.. 2020-04-12. BIP4COVID19: Releasing impact metrics data for articles relevant to COVID-19. https://doi.org/10.1101/2020.04.11.037093

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Dynamics of the COVID -19 Related Publications

BackgroundThis study aims to analyze the dynamics of the published articles and preprints of Covid-19 related literature from different scientific databases and sharing platforms. MethodsThe PubMed, Elsevier, and Research Gate (RG) databases were under consideration in this study over a specific time. Analyses were carried out on the number of publications as (a) function of time (day), (b) journals and (c) authors. Doubling time of the number of publications was analyzed for PubMed "all articles" and Elsevier published articles. Analyzed databases were (1A) PubMed "all articles" (01/12/2019-12/06/2020) (1B) PubMed Review articles (01/12/2019-2/5/2020) and (1C) PubMed Clinical Trials (01/01/2020-30/06/2020) (2) Elsevier all publications (01/12/2019-25/05/2020) (3) RG (Article, Pre Print, Technical Report) (15/04/2020-30/4/2020). FindingsTotal publications in the observation period for PubMed, Elsevier, and RG were 23000, 5898 and 5393 respectively. The average number of publications/day for PubMed, Elsevier and RG were 70.0 {+/-}128.6, 77.6{+/-}125.3 and 255.6{+/-}205.8 respectively. PubMed shows an avalanche in the number of publication around May 10, number of publications jumped from 6.0{+/-}8.4/day to 282.5{+/-}110.3/day. The average doubling time for PubMed, Elsevier, and RG was 10.3{+/-}4 days, 20.6 days, and 2.3{+/-}2.0 days respectively. In PubMed average articles/journal was 5.2{+/-}10.3 and top 20 authors representing 935 articles are of Chinese descent. The average number of publications per author for PubMed, Elsevier, and RG was 1.2{+/-}1.4, 1.3{+/-}0.9, and 1.1{+/-}0.4 respectively. Subgroup analysis, PubMed review articles mean and median review time for each article were <0|17{+/-}17|77> and 13.9 days respectively; and reducing at a rate of-0.21 days (count)/day. InterpretationAlthough the disease has been known for around 6 months, the number of publications related to the Covid-19 until now is huge and growing very fast with time. It is essential to rationalize the publications scientifically by the researchers, authors, reviewers, and publishing houses. FundingNone

scientific communication and education

From bioinformatics user to bioinformatics engineer: a report

Teaching computer programming is not a simple task and it is challenging to introduce the concepts of programming in graduate programs of other fields. Little efforts have been made on engaging students in computational development after programming trainings. An emerging need is to establish subjects of bioinformatics and programming languages in genetics and molecular biology graduate programs, when students in these degree programs are immersed in a sea of genomic and transcriptomic data, which demands proficient computational treatment. I report an empirical guideline to introduce programming languages and recommend Python as first language for graduate programs in which students were from genetics and molecular biology backgrounds. Including the development of programming solutions related to graduate students' research activities may improve programming skills and better engagement. These results suggest that the applied approach leads to enhanced learning of introductory to autonomy in highly advanced programming concepts by graduate students. This guide should be extended for other research programs.

scientific communication and education