bioRxiv Science⌕ Search

Biology subjects

Vora, J.

Publications and source records attributed to Vora, J..

2 recordsLinked to original sources

Enhancing the interoperability of glycan data flow between ChEBI, PubChem, and GlyGen.

Glycans play a vital role in health, disease, bioenergy, biomaterials, and biotherapeutics. As a result, there is keen interest to identify and increase glycan data in bioinformatics databases like ChEBI and PubChem, and connecting them to resources at the EMBL-EBI and NCBI to facilitate access to important annotations at a global level. GlyTouCan is a comprehensive archival database that contains glycans obtained primarily through batch upload from glycan repositories, glycoprotein databases, and individual laboratories. In many instances, the glycan structures deposited in GlyTouCan may not be fully defined or have supporting experimental evidence and citations. Databases like ChEBI and PubChem were designed to accommodate complete atomistic structures with well-defined chemical linkages. As a result, they cannot easily accommodate the structural ambiguity inherent in glycan databases. Consequently, there is a need to improve the organization of glycan data coherently to enhance connectivity across the major NCBI, EMBL-EBI, and glycoscience databases. This paper outlines a workflow developed in collaboration between GlyGen, ChEBI, and PubChem to improve the visibility and connectivity of glycan data across these resources. GlyGen hosts a subset of glycans (~29,000) from the GlyTouCan database and has submitted valuable glycan annotations to the PubChem database and integrated over 10,500 (including ambiguously defined) glycans into the ChEBI database. The integrated glycans were prioritized based on links to PubChem and connectivity to glycoprotein data. The pipeline provides a blueprint for how glycan data can be harmonized between different resources. The current PubChem, ChEBI, and GlyTouCan mappings can be downloaded from GlyGen (https://data.glygen.org).

bioinformatics↗

Differential expression of glycosyltransferases identified through comprehensive pan-cancer analysis

Despite accumulating evidence supporting a role for glycosylation in cancer progression and prognosis, the complexity of the human glycome and glycoproteome poses many challenges to understanding glycosylation-related events in cancer. In this study, a multifaceted genomics approach was applied to analyze the impact of differential expression of glycosyltransferases (GTs) in 16 cancers. An enzyme list was compiled and curated from numerous resources to create a consensus set of GTs. Resulting enzymes were analyzed for differential expression in cancer, and findings were integrated with experimental evidence from other analyses, including: similarity of healthy expression patterns across orthologous genes, miRNA expression, automatically-mined literature, curation of known cancer biomarkers, N-glycosylation impact, and survival analysis. The resulting list of GTs comprises 222 human enzymes based on annotations from five databases, 84 of which were differentially expressed in more than five cancers, and 14 of which were observed with the same direction of expression change across all implicated cancers. 25 high-value GT candidates were identified by cross-referencing multimodal analysis results, including PYGM, FUT6 and additional fucosyltransferases, several UDP-glucuronosyltransferases, and others, and are suggested for prioritization in future cancer biomarker studies. Relevant findings are available through OncoMX at https://data.oncomx.org, and the overarching pipeline can be used as a framework for similarly analysis across diverse evidence types in cancer. This work is expected to improve the understanding of glycosylation in cancer by transparently defining the space of glycosyltransferase enzymes and harmonizing variable experimental data to enable improved generation of data-driven cancer biomarker hypotheses.

genomics↗