bioRxiv · 10.64898/2026.09.12.751192
OpenLipid: a large language model workflow for targeted analysis of DIA mass spectrometry data in lipidomics
Abstract
Mass spectrometry has become a central technology for lipidomics, with data-independent acquisition (DIA) enabling broad and reproducible sampling of lipid signals. However, the multiplexed fragment-ion spectra in DIA data complicate lipid identification. Here, we introduce OpenLipid, a large language model (LLM)-based workflow for targeted DIA lipidomics. Using assay libraries built from data-dependent acquisition (DDA) results, OpenLipid directly evaluates extracted ion chromatograms (XICs) from DIA data in a zero-shot setting to identify target lipid peaks and generate human-readable rationales for individual lipid identification decisions. We benchmarked OpenLipid against manual annotations across four datasets comprising human plasma and mouse feces analyzed in positive and negative ionization modes. The plasma assay libraries contained 199 target lipids in positive mode and 147 in negative mode. The fecal assay libraries contained 181 target lipids in positive mode and 264 in negative mode. At a 5% false discovery rate (FDR) threshold, OpenLipid identified 110 (55.3%) and 28 (19.0%) library targets in plasma and 130 (71.8%) and 84 (31.8%) in feces, in positive and negative ionization modes, respectively. OpenLipid achieved an overall identification rate comparable to that of DIAMetAlyzer (57.8%, 19.7%, 89.0%, and 17.8% across the corresponding datasets) and substantially higher than that of untargeted MS-DIAL DIA analysis (12.6%, 0.0%, 47.5%, and 0.0%) on the same assay-library targets. LLM-derived chromatographic features also enabled supervised discrimination between correct and incorrect candidate peak groups for target lipids, with median cross-validation average precision values of 0.838-0.912. Together, these results demonstrate that OpenLipid is an effective LLM-based workflow for FDR-controlled targeted analysis of DIA lipidomics data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Li, J., Rost, H.. 2026-09-18. OpenLipid: a large language model workflow for targeted analysis of DIA mass spectrometry data in lipidomics. https://doi.org/10.64898/2026.09.12.751192
Cite the original work for its findings. Save a collection to share your selection of sources.