bioRxiv Science⌕ Search

Biology subjects

Colby, S. M.

Publications and source records attributed to Colby, S. M..

3 recordsLinked to original sources

Introducing identification probability for automated and transferable assessment of metabolite identification confidence in metabolomics and related studies

Methods for assessing compound identification confidence in metabolomics and related studies have been debated and actively researched for the past two decades. The earliest effort in 2007 focused primarily on mass spectrometry and nuclear magnetic resonance spectroscopy and resulted in four recommended levels of metabolite identification confidence - the Metabolite Standards Initiative (MSI) Levels. In 2014, the original MSI Levels were expanded to five levels (including two sublevels) to facilitate communication of compound identification confidence in high resolution mass spectrometry studies. Further refinement in identification levels have occurred, for example to accommodate use of ion mobility spectrometry in metabolomics workflows, and alternate approaches to communicate compound identification confidence also have been developed based on identification points schema. However, neither qualitative levels of identification confidence nor quantitative scoring systems address the degree of ambiguity in compound identifications in context of the chemical space being considered, are easily automated, or are transferable between analytical platforms. In this perspective, we propose that the metabolomics and related communities consider identification probability as an approach for automated and transferable assessment of compound identification and ambiguity in metabolomics and related studies. Identification probability is defined simply as 1/N, where N is the number of compounds in a reference library or chemical space that match to an experimentally measured molecule within user-defined measurement precision(s), for example mass measurement or retention time accuracy, etc. We demonstrate the utility of identification probability in an in silico analysis of multi-property reference libraries constructed from the Human Metabolome Database and computational property predictions, provide guidance to the community in transparent implementation of the concept, and invite the community to further evaluate this concept in parallel with their current preferred methods for assessing metabolite identification confidence.

biochemistry↗

Introducing Molecular Hypernetworks forDiscovery in MultidimensionalMetabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and help address the challenge of complete annotation of molecules in untargeted metabolomics. "Molecular networks" (MNs), as used, for example, in the Global Natural Products Social Molecular Networking platform, are an increasingly popular computational strategy for exploring and visualizing molecular relationships and improving annotation. MNs use graph representations to show the re-lationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated to a single molecular identity. How-ever, more advanced methods may better represent the complexity present in samples. Our work aims to increase confidence in annotation propagation by extending molecular network methods to include "molecular hypernetworks" (MHNs), able to natively repre-sent multiway relationships among observations supporting both human and analytical processing. In this paper we first introduce MHNs illustrated with simple examples, and demonstrate how to build them from liquid chromatography-and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs di-rectly from existing MNs as their "clique reconstructions", demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

bioinformatics↗

The unknown lipids project: harmonized methods improve compound identification and data reproducibility in an inter-laboratory untargeted lipidomics study

Untargeted lipidomics allows analysis of a broader range of lipids than targeted methods and permits discovery of unknown compounds. Previous ring trials have evaluated the reproducibility of targeted lipidomics methods, but inter-laboratory comparison of compound identification and unknown feature detection in untargeted lipidomics has not been attempted. To address this gap, five laboratories analyzed a set of mammalian tissue and biofluid reference samples using both their own untargeted lipidomics procedures and a common chromatographic and data analysis method. While both methods yielded informative data, the common method improved chromatographic reproducibility and resulted in detection of more shared features between labs. Spectral search against the LipidBlast in silico library enabled identification of over 2,000 unique lipids. Further examination of LC-MS/MS and ion mobility data, aided by hybrid search and spectral networking analysis, revealed spectral and chromatographic patterns useful for classification of unknown features, a subset of which were highly reproducible between labs. Overall, our method offers enhanced compound identification performance compared to targeted lipidomics, demonstrates the potential of harmonized methods to improve inter-site reproducibility for quantitation and feature alignment, and can serve as a reference to aid future annotation of untargeted lipidomics data.

systems biology↗