Comprehensive LC-MS/MS Data Acquisition in Metabolomics via Maximum Bipartite Matching
BackgroundIn untargeted metabolomics studies, liquid chromatography tandem mass spectrometry (LC-MS/MS) is a powerful analytical platform. The fragmentation spectra produced can be used as "molecular fingerprints" to identify unknown metabolites. However, the high number of analytes that may be co-eluting limits the number of fragmentation spectra that can be collected and potentially identified, presenting a serious bottleneck for many studies. There is a need for new fragmentation strategies which are comprehensive, interpretable and robust, meaning they produce high-quality fragmentation spectra for as many analytes as possible while operating within the constraints of notoriously noisy mass spectrometry data. ResultsWe present a data acquisition workflow which uses a bipartite graph to represent the relationship between opportunities for fragmentation and desired fragmentation targets. This method allows a schedule for data acquisition to be optimally allocated by a standard algorithm. We augment this existing technique by allowing it to solve for multiple samples collectively, allowing it to optimise target intensity (and hence spectral quality) via the use of a weighted matching and by assigning leftover scans redundantly to improve robustness. We also show how this workflow can be used flexibly to generate inclusion windows for Data-Dependent Acquisition (DDA) methods. Our experiments show that several thousand peaks identified in a realistic biological sample can be targeted using only two LC-MS/MS runs. We also further investigate the trade-off between offline workflows and DDA methods by exposing our target list of peaks to realistic variation across samples. We find in those circumstances that our new method has performance (measured by number of peaks targeted comparable to state-of-the-art DDA methods). However, this competitive performance is only possible with our additions to the base maximum matching technique, which provide extra resistance against inter-sample variations. ConclusionsWe have proposed a workflow for LC-MS/MS data acquisition which can be used flexibly for entirely pre-scheduled acquisition or which may generate inclusion windows for online DDA methods. Our results show that the maximum matching workflow with our improvements is state-of-the-art where pre-scheduling is concerned, and in future this foundation may be developed to build more powerful DDA methods which can action the promise of truly comprehensive data acquisition.