bioRxiv Science⌕ Search

Biology subjects

Bolton, E.

Publications and source records attributed to Bolton, E..

2 recordsLinked to original sources

Breaking Through Biology's Data Wall: Expanding the Known Tree of Life by Over 10x using a Global Biodiscovery Pipeline

Advancements in the life sciences have always been built upon our collective understanding of life on Earth. Now, the rise of generative biology - the use of AI foundation models to design, generate, and annotate proteins, pathways and therapeutics - is creating unprecedented demand for large, diverse biological sequence datasets. While a limited subset of such data can be generated in clinical or laboratory settings, the vast majority of the training data for unsupervised models must be sourced from the natural world - the product of nearly four billion years of evolutionary history. However, the public databases that currently supply this data, while foundational to research, were established to aggregate results from academic experiments, not as training datasets for machine learning. Their human-centric data structure limits model performance due to redundancy, taxonomic and geographic bias, limited biological context, and inconsistent provenance. With 68% of all sequence data in the SRA database coming from just 5 species, this is one of the most severe class imbalance problems ever encountered in AI. Legal and infrastructural constraints further exacerbate this bottleneck. To address these limitations and support scalable model training, we introduce BaseData: the largest and fastest-growing biological sequence database ever built, and the first purpose-built for training foundation models. As of late 2024, BaseData contained 9.8 billion novel genes, representing more than a 10-fold expansion in known protein diversity after accounting for redundancy. BaseData also contains more than 1 million species not represented in other genomic databases. Its partnership-driven data supply chain across 26 countries and autonomous regions enables growth of up to 2 billion novel genes per month, far exceeding public repositories. All data is collected under benefit sharing agreements using standardized protocols and structured using graph-based, ontology-rich metadata that preserves evolutionary context. BaseData represents a new, ethically-grounded infrastructure for training biological foundation models, complementing public efforts and enabling the next era of generative biology. Short AbstractProgress of AI in biology is now being limited by the availability of high-quality biological sequence data from nature. To get past this data wall, we introduce BaseData, a new biological database built on top of a global, scalable biodiscovery pipeline. As of 2024, BaseData had already expanded the known protein universe by over 10x.

genomics↗

Enhancing the interoperability of glycan data flow between ChEBI, PubChem, and GlyGen.

Glycans play a vital role in health, disease, bioenergy, biomaterials, and biotherapeutics. As a result, there is keen interest to identify and increase glycan data in bioinformatics databases like ChEBI and PubChem, and connecting them to resources at the EMBL-EBI and NCBI to facilitate access to important annotations at a global level. GlyTouCan is a comprehensive archival database that contains glycans obtained primarily through batch upload from glycan repositories, glycoprotein databases, and individual laboratories. In many instances, the glycan structures deposited in GlyTouCan may not be fully defined or have supporting experimental evidence and citations. Databases like ChEBI and PubChem were designed to accommodate complete atomistic structures with well-defined chemical linkages. As a result, they cannot easily accommodate the structural ambiguity inherent in glycan databases. Consequently, there is a need to improve the organization of glycan data coherently to enhance connectivity across the major NCBI, EMBL-EBI, and glycoscience databases. This paper outlines a workflow developed in collaboration between GlyGen, ChEBI, and PubChem to improve the visibility and connectivity of glycan data across these resources. GlyGen hosts a subset of glycans (~29,000) from the GlyTouCan database and has submitted valuable glycan annotations to the PubChem database and integrated over 10,500 (including ambiguously defined) glycans into the ChEBI database. The integrated glycans were prioritized based on links to PubChem and connectivity to glycoprotein data. The pipeline provides a blueprint for how glycan data can be harmonized between different resources. The current PubChem, ChEBI, and GlyTouCan mappings can be downloaded from GlyGen (https://data.glygen.org).

bioinformatics↗