bioRxiv Science⌕ Search

Biology subjects

Devreese, R.

Publications and source records attributed to Devreese, R..

4 recordsLinked to original sources

DeepLC introduces transfer learning for accurate LC retention time prediction and adaptation to substantially different modifications and setups

While LC retention time prediction of peptides and their modifications has proven useful, widespread adoption and optimal performance are hindered by variations in experimental parameters. These variations can render retention time prediction models inaccurate and dramatically reduce the value of predictions for identification, validation, and DIA spectral library generation. To date, mitigation of these issues has been attempted through calibration or by training bespoke models for specific experimental setups, with only partial success. We here demonstrate that transfer learning can successfully overcome these limitations by leveraging pre-trained model parameters. Remarkably, this approach can even fit highly performant models to substantially different peptide modifications and LC conditions than those on which the model was originally trained. This impressive adaptability of transfer learning makes it a highly robust solution for accurate peptide retention time prediction across a very wide variety of imaginable proteomics workflows.

bioinformatics↗

Collisional cross-section prediction for multiconformational peptide ions with IM2Deep

Peptide collisional cross-section (CCS) prediction is complicated by the tendency of peptide ions to exhibit multiple conformations in the gas phase. This adds further complexity to downstream analysis of proteomics data, for example for identification or quantification through feature finding. Here, we present an improved version of IM2Deep that is trained on a carefully curated dataset to predict CCS values of multiconformational peptides. The training data is derived from a large and comprehensive set of publicly available datasets. This comprehensive training dataset together with a tailored architecture allows for the accurate CCS prediction of multiple peptide conformational states. Furthermore, the enhanced IM2Deep model also retains high precision for peptides with a single observed conformation. IM2Deep is publicly available under a permissive open source license at https://github.com/compomics/IM2Deep.

bioinformatics↗

Maximizing immunopeptidomics-based bacterial epitope discovery by multiple search engines and rescoring

Mass spectrometry-based discovery of bacterial immunopeptides presented by infected cells allows untargeted discovery of bacterial antigens that can serve as vaccine candidates. However, reliable identification of bacterial epitopes is challenged by their extreme low abundance. Here, we describe an optimized bioinformatical framework to enhance the confident identification of bacterial immunopeptides. Immunopeptidomics data of cell cultures infected with Listeria monocytogenes were searched by four different search engines, PEAKS, Comet, Sage and MSFragger, followed by data-driven rescoring with MS2Rescore. Compared to individual search engine results, this integrated workflow boosted immunopeptide identification by an average of 27% and led to the high-confidence detection of 18 additional bacterial peptides (+27%) matching 15 different Listeria proteins (+36%). Despite the strong agreement between the search engines, a small number of spectra (< 1%) had ambiguous matches to multiple peptides and were excluded to ensure high-confident identifications. Finally, we demonstrate our workflow with sensitive timsTOF SCP data acquisition and find that rescoring, now with inclusion of ion mobility features, identifies 76% more peptides compared to Q Exactive HF acquisition. Together, our results demonstrate how integration of multiple search engine results along with data-driven rescoring maximizes immunopeptide identification, boosting the detection of high-confidence bacterial epitopes for vaccine development.

molecular biology↗

TIMS2Rescore: A DDA-PASEF optimized data-driven rescoring pipeline based on MS2Rescore

The high throughput analysis of proteins with mass spectrometry (MS) is highly valuable for understanding human biology, discovering disease biomarkers, identifying therapeutic targets, and exploring pathogen interactions. To achieve these goals, specialized proteomics subfields - such as plasma proteomics, immunopeptidomics, and metaproteomics - must tackle specific analytical challenges, such as an increased identification ambiguity compared to routine proteomics experiments. Technical advancements in MS instrumentation can counter these issues by acquiring more discerning information at higher sensitivity levels, as is exemplified by the incorporation of ion mobility and parallel accumulation - serial fragmentation (PASEF) technologies in timsTOF instruments. In addition, AI-based bioinformatics solutions can help overcome ambiguity issues by integrating more data into the identification workflow. Here, we introduce TIMS2Rescore, a data-driven rescoring workflow optimized for DDA-PASEF data from timsTOF instruments. This platform includes new timsTOF MS2PIP spectrum prediction models and IM2Deep, a new deep learning-based peptide ion mobility predictor. Furthermore, to fully streamline data throughput, TIMS2Rescore directly accepts Bruker raw mass spectrometry data, and search results from ProteoScape and many other search engines, including MS Amanda and PEAKS. We showcase TIMS2Rescore performance on plasma proteomics, immunopeptidomics (HLA class I and II), and metaproteomics data sets. TIMS2Rescore is open-source and freely available at https://github.com/compomics/tims2rescore.

bioinformatics↗