bioRxiv Science⌕ Search

Biology subjects

Koblitz, J.

Publications and source records attributed to Koblitz, J..

2 recordsLinked to original sources

The StrainDiscoveryDatabase: an open framework for standardized microbial strain data

The vast amount of existing data on microbial strains holds immense potential to revolutionize bioindustry through the application of Artificial Intelligence (AI). However, the training of robust predictive AI models requires large-scale, unified, and non-redundant microbial datasets, which is currently severely hindered by the deep fragmentation of the data and the existence of synonymous strain identifiers in different culture collections. To overcome these infrastructural bottlenecks, we have established the StrainDiscoveryDatabase (SDD), a comprehensive, machine-readable dataset encompassing over 6.2 million harmonized data points for 256,889 microbial strains. The SDD does not rely on its own data repository, but rather on existing data that is retrieved on the fly from highly curated databases. Through an automated pipeline phenotypic, genotypic, and contextual data are systematically retrieved via the Application Programming Interfaces (APIs) of the Bacterial Diversity database (BacDive), the Microbial Resource Research Infrastructure Information System (MIRRI-IS) and the catalogue of the DSMZ. In order to reliably resolve synonymous strain identifiers, the StrainInfo database and its identification tools are employed, enabling the accurate deduplication and unification of records from disparate sources. The resulting aggregated, globally unique dataset is provided in a highly standardized JSON format in strict adherence to the FAIR data principles. By bridging isolated database silos and linking distributed knowledge to discrete biological entities, the SDD provides a high-quality, foundational resource designed to accelerate trait-based strain discovery, large-scale comparative analysis, and machine learning applications, thereby supporting the translation of the extensive existing knowledge on microbial traits to bioindustrial applications.

microbiology↗

Predicting bacterial phenotypic traits through improved machine learning using high-quality, curated datasets

Predicting prokaryotic phenotypes - observable traits that govern functionality, adaptability, and interactions - holds significant potential for fields such as biotechnology, environmental sciences, and evolutionary biology. This study leverages machine learning to explore the relationship between prokaryotic genotypes and phenotypes. Taking advantage of the highly standardized datasets in the BacDive database, we modeled eight physiological properties based on protein family inventories, discuss the evaluation metrics, and explore the biological implications of our models. The high confidence values of our predictions highlight the importance of data quality and quantity for a reliable inference of bacterial phenotypes. Our approach yielded nearly 55,000 new data points for approximately 20,000 strains which are published openly in the BacDive database, enriching existing phenotypic datasets and paving the way for future research and analysis. The open-source software generated can readily be applied to other datasets, for example the IMG/M system for metagenomics, as well as different applications, like the assessment of the potential of soil bacteria for bioremediation projects.

bioinformatics↗