bioRxiv Science⌕ Search

Biology subjects

Wilding-McBride, D.

Publications and source records attributed to Wilding-McBride, D..

3 recordsLinked to original sources

A de novo MS1 feature detector for the Bruker timsTOF Pro

1Identification of peptides by analysis of data acquired by the two established methods for bottom-up proteomics, DDA and DIA, relies heavily on the fragment spectra. In DDA, peptide features detected in mass spectrometry data are identified by matching their fragment spectra with a peptide database. In DIA, a peptides fragment spectra are targeted for extraction and matched with observed spectra. Although fragment ion matching is a central aspect in most peptide identification strategies, the precursor ion in the MS1 data reveals important characteristics as well, including charge state, intensity, monoisotopic m/z, and apex in retention time. Most importantly, the precursors mass is essential in determining the potential chemical modification state of the underlying peptide sequence. In the timsTOF, with its additional dimension of collisional cross-section, the data representing the precursor ion also reveals the peptides peak in ion mobility. However, the availability of tools to survey precursor ions with a wide range of abundance in timsTOF data across the full mass range is very limited. Here we present a de novo feature detector called three-dimensional intensity descent (3DID). 3DID can detect and extract peptide features down to a configurable intensity level, and finds many more features than several existing tools. 3DID is written in Python and is freely available with an open-source MIT license to facilitate experimentation and further improvement (DOI 10.5281/zenodo.6513126). The dataset used for validation of the algorithm is publicly available (ProteomeXchange identifier PXD030706). 2 Author SummaryIn the identification of peptides in mass spectrometry data, much attention has been given to the targeting and extraction of mass spectra produced by fragmentation of precursor ions. However, important information about the peptide is revealed by the data representing the precursor ion itself, such as the peptides charge state, mass-to-charge ratio, intensity, and retention time. The timsTOF produces the additional dimension of ion mobility, which provides richer information about the precursor. Although tools exist for the analysis of timsTOF data, they are hampered by limited dynamic range. In this work, we describe a de novo feature detector called 3DID that detects peptide features across the full mass range. Our detector can detect more peptides than existing tools across a broader range of abundance, which enables more comprehensive analysis of the data. We believe 3DID will make a valuable contribution to the proteomics toolbox.

bioinformatics↗

Predicting coordinates of peptide features in raw timsTOF data with machine learning for targeted extraction reduces missing values in label-free DDA LC-MS/MS proteomics experiments

1The determination of relative protein abundance in label-free data dependant acquisition (DDA) LC-MS/MS proteomics experiments is hindered by the stochastic nature of peptide detection and identification. Peptides with an abundance near the limit of detection are particularly effected. The possible causes of missing values are numerous, including; sample preparation, variation in sample composition and the corresponding matrix effects, instrument and analysis software settings, instrument and LC variability, and the tolerances used for database searching. There have been many approaches proposed to computationally address the missing values problem, predominantly based on transferring identifications from one run to another by data realignment, as in MaxQuants matching between runs (MBR) method, and/or statistical imputation. Imputation transfers identifications by statistical estimation of the likelihood the peptide is present based on its presence in other technical replicates but without probing the raw data for evidence. Here we present a targeted extraction approach to resolving missing values without modifying or realigning the raw data. Our method, which forms part of an end-to-end timsTOF processing pipeline we developed called Targeted Feature Detection and Extraction (TFD/E), predicts the coordinates of peptides using machine learning models that learn the delta of each peptides coordinates from a reference library. The models learn the variability of a peptides location in 3D space from the variability of known peptide locations around it. Rather than realigning or altering the raw data, we create a run-specific lens through which to observe the data, targeting a location for each peptide of interest and extracting it. By also creating a method for extracting decoys, we can estimate the false discovery rate (FDR). Our method outperforms MaxQuant and MSFragger by achieving substantially fewer missing values across an experiment of technical replicates. The software has been developed in Python using Numpy and Pandas and open sourced with an MIT license (DOI 10.5281/zenodo.6513126) to provide the opportunity for further improvement and experimentation by the community. Data are available via ProteomeXchange with identifier PXD030706. 2 Author SummaryMissed identifications of peptides in data-dependent acquisition (DDA) proteomics experiments are an obstacle to the precise determination of which proteins are present in a sample and their relative abundance. Efforts to address the problem in popular analysis workflows include realigning the raw data to transfer a peptide identification from one run to another. Another approach is statistically analysing peptide identifications across an experiment to impute peptide identifications in runs in which they were missing. We propose a targeted extraction technique that uses machine learning models to construct a run-specific lens through which to examine the raw data and predict the coordinates of a peptide in a run. The models are trained on differences between observations of confidently identified peptides in a run and a reference library of peptide observations collated from multiple experiments. To minimise the risk of drawing unsound experimental conclusions based on an unknown rate of false discoveries, our method provides a mechanism for estimating the false discovery rate (FDR) based on the misclassification of decoys as target features. Our approach outperforms the popular analysis tool suites MaxQuant and MSFragger/IonQuant, and we believe it will be a valuable contribution to the proteomics toolbox for protein quantification.

bioinformatics↗

Simplifying MS1 and MS2 spectra to achieve lower mass error, more dynamic range, and higher peptide identification confidence on the Bruker timsTOF Pro

1For bottom-up proteomic analysis, the goal of analytical pipelines that process the raw output of mass spectrometers is to detect, characterise, identify, and quantify peptides. The initial steps of detecting and characterising features in raw data must overcome some considerable challenges. The data presents as a sparse array, sometimes containing billions of intensity readings over time. These points represent both signal and chemical or electrical noise. Depending on the biological samples complexity, tens to hundreds of thousands of peptides may be present in this vast data landscape. For ion mobility-based LC-MS analysis, each peptide is comprised of a grouping of hundreds of single intensity readings in three dimensions: mass-over-charge (m/z), mobility, and retention time. There is no inherent information about any associations between individual points; whether they represent a peptide or noise must be inferred from their structure. Peptides each have multiple isotopes, different charge states, and a dynamic range of intensity of over six orders of magnitude. Due to the high complexity of most biological samples, peptides often overlap in time and mobility, making it very difficult to tease apart isotopic peaks, to apportion the intensity of each and the contribution of each isotope to the determination of the peptides monoisotopic mass, which is critical for the peptides identification. Here we describe four algorithms for the Bruker timsTOF Pro that each play an important role in finding peptide features and determining their characteristics. These algorithms focus on separate characteristics that determine how candidate features are detected in the raw data. The first two algorithms deal with the complexity of the raw data, rapidly clustering raw data into spectra that allows isotopic peaks to be resolved. The third algorithm compensates for saturation of the instruments detector thereby recovering lost dynamic range, and lastly, the fourth algorithm increases confidence of peptide identifications by simplification of the fragment spectra. These algorithms are effective in processing raw data to detect features and extracting the attributes required for peptide identification, and make an important contribution to an analytical pipeline by detecting features that are higher quality and better segmented from other peptides in close proximity. The software has been developed in Python using Numpy and Pandas and made freely available with an open-source MIT license to facilitate experimentation and further improvement (DOI 10.5281/zenodo.6513126). Data are available via ProteomeXchange with identifier PXD030706. 2 Author SummaryThe primary goal of mass spectrometry data processing pipelines in the proteomic analysis of complex biological samples is to identify peptides accurately and comprehensively with abundance across a broad dynamic range. It has been reported that detection of low-abundance peptides for early-disease biomarkers in complex fluids is limited by the sensitivity of biomarker discovery platforms (1), the dynamic range of plasma abundance, which can exceed ten orders of magnitude (2), and the fact that lower abundance proteins provide the most insight in disease processes (3). As mass spectrometry hardware improves, the corresponding increase in amounts of data for analysis pushes legacy software analysis methods out of their designed specification. Additionally, experimentation with new algorithms to analyse raw data produced by instruments such as the Bruker timsTOF Pro has been hampered by the paucity of modular, open-source software pipelines written in languages accessible by the large community of data scientists. Here we present several algorithms for simplifying MS1 and MS2 spectra that are written in Python. We show that these algorithms are effective to help improve the quality and accuracy of peptide identifications.

bioinformatics↗