bioRxiv Science⌕ Search

Biology subjects

Golomb Duran, T.

Publications and source records attributed to Golomb Duran, T..

2 recordsLinked to original sources

Specifind: A Natural Language Processing Tool for Automating Species Occurrence (Re-)Discovery from Scientific Literature

A vast amount of valuable information on species occurrences remains embedded within the unstructured and continuously expanding body of ecological literature written in natural language. When effectively extracted, such dispersed knowledge has the potential to significantly improve understanding of species distributions and ecological patterns, thereby enabling more targeted and informed conservation actions. To address this challenge, Specifind is introduced as a tool designed to facilitate the extraction of species occurrence data through the identification of scientific species names, geographic references, and the relationships between them while ensuring traceability. At its core lies a newly developed and expertly annotated dataset comprising over one thousand open-access abstracts drawn from five domains: biogeography, botany, entomology, mycology, and zoology. A suite of Natural Language Processing components has been integrated, including Document Layout Analysis, Optical Character Recognition, Named Entity Recognition, Coreference Resolution, and Relation Extraction. These components enable the accurate identification and contextual linking of species and geographic entities, even when references are distributed throughout the text. As a result, Specifind enhances the discoverability and usability of species occurrence information embedded within unstructured scientific texts. It is expected to reduce the substantial effort required for manual literature review, thereby supporting biodiversity research, informing conservation planning, and facilitating ecological analysis.

bioinformatics↗

biodumpy: A Comprehensive Biological Data Downloader

In recent years, the expansion of public biodiversity platforms and associated datasets has greatly improved access to ecological and biological information. These resources now cover vast geographic areas, extended temporal scales, and diverse taxonomic groups, becoming essential for ecological studies by enabling more comprehensive analyses and novel hypotheses testing. Concurrently, the development of programming packages has facilitated data access and interaction, streamlining their retrieval processes. However, most existing tools are limited to specific databases, posing challenges for studies requiring seamless integration of data from multiple sources. The growing availability of biodiversity data highlights the urgent need for robust tools to efficiently process, analyse, and interpret ecological and biological information. To address this limitation, we introduce biodumpy, a new Python package developed for the retrieval, management, and integration of biological data from various public databases. biodumpy provides access to up-to-date and comprehensive datasets spanning genetic, distributional, taxonomic, and bibliographic sources. It includes specialized modules for efficient data retrieval across taxonomic lists, with the capability to process multiple modules simultaneously. By integrating diverse data sources, biodumpy enhances data acquisition, providing researchers with a powerful framework for comprehensive analyses and supporting ecological research to tackle complex environmental challenges.

ecology↗