bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.07.30.741704

Quantifying per-match Reliability in Library Matching for Untargeted Metabolomics Workflows

Abstract

Tandem mass spectrometry has become central to untargeted metabolomics. The translation of unknown spectra into biological insight depends on assigning chemical identities to detected metabolites. Structural characterization typically begins with mass spectral library matching, in which experimental spectra are compared against reference libraries and candidate annotations are ranked by their spectral similarity to the query. As spectral libraries and experimental datasets grow, however, more candidates achieve comparable similarity scores for a single query, and similarity scores give no indication of how reproducible a candidate match is or how sensitive it is to the underlying fragment evidence. Existing false-discovery-rate approaches can indicate annotation error at the dataset level but do not provide a per-match estimate of reliability. Here, we introduce a SpecReBoot-inspired query-focused bootstrapping approach that resamples the fragment evidence of each query spectrum. This approach relies on recomputing query similarity to candidate library spectra across bootstrap replicates, which provides a statistical distribution of scores rather than a single value. From this distribution we define the match support, a per-match reliability estimate quantifying the reproducibility of a match under spectral perturbation, together with measures of ranking stability that describe how often a candidate remains among the top-ranked matches across replicates. Applied to a forensic drug-of-abuse case, match support distinguished previously identified annotations from high-scoring false positives: a distinction cosine similarity failed to make. Furthermore, match support values remained stable as the reference library was expanded, whereas ranking stability metrics shifted significantly. In a cross-instrument endogenous metabolite library search, match support further revealed metric-specific annotation behavior, identifying metabolites consistently supported across different similarity metrics, while flagging annotations whose reliability depended strongly on the chosen scoring metric. Benchmarking against a natural-product reference library demonstrated that ranking based on match support values promoted true matches by four ranks on average compared with cosine-based ranking, without promoting analogs. Under controlled spectral perturbation experiments, match support flagged incorrect annotations with an AUROC of 0.75, whereas the cosine similarity score alone of the same match reached only 0.56. Query-focused bootstrapping thus provides a practical, per-match measure of annotation reliability, bringing the field a step toward reliable annotations at scale. We anticipate that incorporation of our annotation reliability scoring into computational metabolomics workflows will further promote the growth of spectral libraries and enhance their applicability across scientific disciplines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Charria-Giron, E., van IJcken, J., Della Vedova, L., Torres-Ortega, L. R., van der Hooft, J. J. J.. 2026-07-30. Quantifying per-match Reliability in Library Matching for Untargeted Metabolomics Workflows. https://doi.org/10.64898/2026.07.30.741704

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

aaRSID, an engineered pyrrolysyl-tRNA synthetase platform for multi-probe proximity proteomics

Proximity labeling (PL) methods utilize spatially targeted chemical or enzymatic generation of a diffusible, reactive intermediate to covalently tag neighboring proteins in living systems. Unlike other tools for studying molecular interactions, PL can detect transient protein relationships with high spatial and temporal sensitivity, allowing for insight into their roles in biological processes. However, current enzymatic PL tools, such as TurboID and APEX2, are limited by their substrate structure and chemistry, which can generate significant background and/or perturb cellular physiology. To address these limitations, we have developed aminoacyl-tRNA synthetase ID (aaRSID), a PL tool that leverages an engineered pyrrolysyl tRNA synthetase (PylRS) for proximity labeling of proteins. We chose PylRS because it can catalyze promiscuous lysine labeling in the absence of its cognate tRNA and utilize a variety of non-canonical amino acids (ncAAs) as substrates. Here, we demonstrate aaRSID's intrinsic proximity labeling activity, use directed evolution to improve this activity, and apply the improved mutant (aaRSID-Ma1.3) for subcellular proteomics and multiplexed imaging. Our work establishes aminoacyl-tRNA synthetases as a new PL enzyme class and introduces a versatile chemical platform for developing ncAA-derived probes to map cellular microenvironments, greatly expanding the applications possible of PL technology.

biochemistry↗

Cellular uptake of folate-olaparib conjugates via folate receptor-mediated endocytosis: Potential for selective delivery of DNA damage response inhibitors into tumour cells

The folate receptor (FR) is overexpressed in a range of human tumours including ovarian cancer cells. We propose that the overexpression of the FR on the surface of ovarian tumour cells could be exploited for the selective delivery of a DNA damage response inhibitor (DDRi) in the form of an intact folate drug conjugate (FDC). This approach would improve the therapeutic index of the parent DDRi facilitating combination studies of the DDRi-based FDC with DNA damaging chemotherapy. FR-mediated cellular uptake of the proposed folate drug conjugates is requisite for FDC selective delivery into tumours. In this study, we synthesised a series of olaparib-based folate conjugates that maintained the biochemical PARP1 inhibition associated with olaparib and showed binding affinity for the folate receptor. Significantly, we identified compounds 10b and 11 that selectively enter FR overexpressing tumour cells via folate receptor-mediated endocytosis in their intact form and engage with their target as demonstrated by the potent inhibition of PARylation (KB cells, PARylation IC50 = 5.7 and 3.9 nM; respectively).

biochemistry↗

Architecture and Energy Transfer of the Bacterial Photosynthetic Unit

In phototrophic organisms, pigment-protein membrane complexes are densely packed to form photosynthetic units (PSUs) that capture solar energy and convert it into chemical energy. Although the structures of many individual photosynthetic complexes have been resolved, how they are arranged and interact with others within photosynthetic membranes to enable efficient excitation energy transfer (EET) remains poorly understood. Here, we report cryo-electron microscopy structures of PSU supercomplex assemblies from the phototrophic a-proteobacterium Rhodovulum viride, including an RC-LH1 core associated with one or two peripheral LH2 complexes and a curved LH2 tetramer. These membrane-derived assemblies define the relative positions and orientations of neighboring photosynthetic complexes and place their pigment arrays in proximity across antenna-antenna and antenna-core interfaces. Structure-based simulations identify potential EET pathways within the PSU assemblies and reveal rapid energy transfer across both LH2-LH2 and LH2-LH1 interfaces. Collectively, these findings provide insights into the assembly and structural modularity of bacterial PSUs and elucidate how the lateral organization of membrane protein complexes facilitates efficient energy transfer. This work extends structural studies of bacterial photosynthesis from individual complexes to their native higher-order assembly, providing a framework for understanding how photosynthetic supercomplex organization shapes energy migration and for guiding the design of artificial photosynthesis.

biochemistry↗