bioRxiv Science⌕ Search

bioRxiv · 10.1101/2021.07.13.452159

PDF Data Extractor (PDE) - A Free Web Application and R Package Allowing the Extraction of Tables from Portable Document Format (PDF) Files and High-Throughput Keyword Searches of Full-Text Articles

Abstract

Literature reviews are generally time-consuming and rely heavily on accurate representation of the data in the title and abstract of articles. Often minor results and details are lost in a systematic screen, which is becoming even more frequent with the rapidly rising numbers of daily published scientific articles. We developed the PDF Data Extractor (PDE) R package to aid scientists at any stage in literature reviews while offering a user-friendly interface. The tool permits the user to categorize large numbers of full-text articles in PDF format, export containing tables to Excel sheets (pdf2table), and extract relevant data using a simple user interface, requiring no bioinformatics skills. Specific features of the literature analysis comprise the adaptability of analysis parameters including the use of regular expressions, machine learning-powered detection of abbreviations of search words in articles, and the export of document meta-data. We exemplify how the PDE R package can be utilized as a pre-screening tool allowing automated categorization of full-text articles by relevance, thereby reducing the literature to be evaluated (in our example by 35% with a sensitivity of 100% at standard parameters). The PDE R package is available from the Comprehensive R Archive Network at https://CRAN.R-project.org/package=PDE and as web tool with limited capacity at https://erikstricker.shinyapps.io/PDE_analyzer/.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stricker, E., Scheurer, M. E.. 2021-07-13. PDF Data Extractor (PDE) - A Free Web Application and R Package Allowing the Extraction of Tables from Portable Document Format (PDF) Files and High-Throughput Keyword Searches of Full-Text Articles. https://doi.org/10.1101/2021.07.13.452159

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

SIMHYB 2: a software tool to explore and illustrate evolutionary forces in Population Genetics teaching and research. Application to Conservation Genetics

Practical approaches have become a standard in many scientific disciplines, including population genetics. By analyzing properly selected datasets, the students can calculate parameters and draw conclusions about genetic diversity, differentiation and evolution of populations with higher efficiency than if based exclusively on theoretical lessons. However, preparing the appropriate datasets is a hard task and a wrong selection can spoil a well-aimed practice. Here we present SO_SCPLOWIMC_SCPLOWHO_SCPLOWYBC_SCPLOW 2, a software tool specifically intended to ease the full understanding of evolutionary forces by the students and to help the teacher to prepare adequate datasets and examples for the practices. It simulates the course of a mixed population under user-defined reproductive and evolutionary conditions. Outputs can be easily adapted for downstream analysis with other popular tools as GO_SCPLOWENC_SCPLOWAO_SCPLOWLC_SCPLOWEO_SCPLOWXC_SCPLOW or SO_SCPLOWTRUCTUREC_SCPLOW. Thus, SO_SCPLOWIMC_SCPLOWHO_SCPLOWYBC_SCPLOW 2 is very suitable for project-based-learning approaches: students can produce their own datasets in different scenarios of genetic drift, migration, selective advantage, reproductive success... Additionally, SO_SCPLOWIMC_SCPLOWHO_SCPLOWYBC_SCPLOW 2 is the only simulation software available to date providing traceable pedigrees of individuals, being therefore very convenient for preparing datasets for parentage analysis, spatial genetic structure or conservation genetics study cases. Satisfactory results from its ongoing utilization in higher education and research are reported.

scientific communication and education↗

Postdoctoral Scholar Recruitment and Hiring Practices in STEM: A Pilot Study

Despite the importance of the postdoctoral position in the training of scientists for independent research careers, few studies have addressed recruiting and hiring of postdocs. We conducted a pilot study on postdoctoral hiring in the Division of Chemistry and Chemical Engineering at the California Institute of Technology to serve as a starting point to better understand postdoctoral recruiting and hiring processes. From this survey of both postdocs and faculty, together with the available literature, the picture emerges that the postdoc hiring process is more decentralized than either faculty hiring or graduate admissions. Postdoc positions are often filled through a passive process where the initial expression of interest from a prospective postdoc is through a "cold-call" contact to a prospective advisor. Individual faculty members are often responsible for developing and implementing their own outreach and recruitment plans and deciding who to hire into a postdoc position. The overall opacity of the processes and practices by which postdocs are identified, recruited, and hired make it difficult to pinpoint where interventions could be effective to ensure equitable hiring practices. Implementation of such practices is critical to training a diverse postdoc population and subsequently of the future STEM faculty recruited from this group.

scientific communication and education↗