bioRxiv Science⌕ Search

Biology subjects

Silverstein, J. C.

Publications and source records attributed to Silverstein, J. C..

2 recordsLinked to original sources

Petagraph: A large-scale unifying knowledge graph framework for integrating biomolecular and biomedical data

The use of biomedical knowledge graphs (BMKG) for knowledge representation and data integration has increased drastically in the past several years due to the size, diversity, and complexity of biomedical datasets and databases. Data extraction from a single dataset or database is usually not particularly challenging. However, if a scientific question must rely on integrative analysis across multiple databases or datasets, it can often take many hours to correctly and reproducibly extract and integrate data towards effective analysis. To overcome this issue, we created Petagraph, a large-scale BMKG that integrates biomolecular data into a schema incorporating the Unified Medical Language System (UMLS). Petagraph is instantiated on the Neo4j graph platform, and to date, has fifteen integrated biomolecular datasets. The majority of the data consists of entities or relationships related to genes, animal models, human phenotypes, drugs, and chemicals. Quantitative data sets containing values from gene expression analyses, chromatin organization, and genetic analyses have also been included. By incorporating models of biomolecular data types, the datasets can be traversed with hundreds of ontologies and controlled vocabularies native to the UMLS, effectively bringing the data to the ontologies. Petagraph allows users to analyze relationships between complex multi-omics data quickly and efficiently.

bioinformatics↗

Tissue Registration and Exploration User Interfaces in support of a Human Reference Atlas

Several international consortia are collaborating to construct a human reference atlas, which is a comprehensive, high-resolution, three-dimensional atlas of all the cells in the healthy human body. Laboratories around the world are collecting tissue specimens from donors varying in sex, age, ethnicity, and body mass index. However, integrating and harmonizing tissue data across 20+ organs and more than 15 bulk and spatial single-cell assay types poses diverse challenges. Here we present the software tools and user interfaces developed to annotate ("register") and explore the collected tissue data. A key part of these tools is a common coordinate framework, which provides standard terminologies and data structures for describing specimens, biological structures, and spatial positions linked to existing ontologies. As of December 2021, the "registration" user interface has been used to harmonize and make publicly available data on 6,178 tissue sections from 2,698 tissue blocks collected by the Human Biomolecular Atlas Program, the Stimulating Peripheral Activity to Relieve Conditions program, the Human Cell Atlas, the Kidney Precision Medicine Project, and the Genotype Tissue Expression project. The second "exploration" user interface enables consortia to evaluate data quality and coverage, explore tissue data in the context of the human body, and guide data acquisition.

bioinformatics↗