bioRxiv Science⌕ Search

Biology subjects

Nameni, A.

Publications and source records attributed to Nameni, A..

6 recordsLinked to original sources

Evaluating the use of non-linear models in data-driven rescoring of peptide-spectrum matches

In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly supports this task through peptide-spectrum match (PSM) rescoring, in which a classifier, typically a linear semi-supervised model, refines the initial matching score. However, Mokapot allows the user to choose among different machine learning algorithms of increasing complexity, from the default linear support vector machine (LSVM) to random forest and XGBoost. Here, we use an entrapment approach to assess the effect of this increasing complexity on PSM identification and the accuracy of the estimated false discovery rate (FDR). We show that, while more complex models increase the number of identified PSMs at a fixed FDR threshold, this gain reflects a bias towards random matches from the target proteome database rather than genuine identifications. Indeed, for the most complex model, the entrapment FDR reaches 6.3% instead of the estimated 1% decoy FDR. This bias thus yields overly optimistic FDR estimates, indicating that model complexity in PSM rescoring must be carefully balanced against this overfitting risk.

bioinformatics↗

Predicting and Elucidating Peptide Retention Mechanisms with Graph Attention Networks

Liquid chromatography (LC) is a key technology in bottom-up proteomics, separating proteolytic peptides to decrease sample complexity, enhance coverage, and increase the robustness of protein identification and quantification. Although high-resolution mass spectrometry has advanced significantly, comparable progress in LC has lagged, primarily due to a limited understanding of peptide-column interactions. To bridge this knowledge gap, we introduce a novel deep learning model (PeptideGNN) based on a Graph Neural Network (GNN) architecture to model and elucidate peptide behaviors across various separation conditions. Trained to accurately predict peptide retention times on ten diverse proteomic datasets, the model subsequently employed a saliency mapping technique to interpret the underlying retention mechanisms. Our model consistently outperformed existing retention-time predictors across multiple datasets, while the saliency mapping, importantly, revealed insights into peptide-stationary phase interactions, highlighting the effects of neighboring amino acids, post-translational modifications (PTMs), chromato-graphic columns, and mobile phase additives on peptide retention.

bioinformatics↗

ProteoBench: the community-curated platform for comparing proteomics data analysis workflows

Mass spectrometry (MS)-based proteomics is a well-established strategy for analyzing complex biological mixtures. Many MS instruments and data acquisition strategies are available, and the data they acquire differ substantially, thus requiring tailored analysis algorithms. Hence, many dedicated bioinformatics workflows are developed. These are in constant evolution, and the community lacks a centralized platform for comparing their performance. Here, we propose ProteoBench, a single platform that brings together software developers and software users to provide an ever-evolving comparison of state-of-the-art proteomics data processing tools. ProteoBench is an open-source resource that enables the community to evaluate data analysis workflows, develop benchmarking modules dedicated to specific comparisons, and discuss the best methods to compare software tools. The platform ensures that the benchmark evolves alongside advances in proteomics data analysis workflows. ProteoBench guides researchers towards the best-suited tool and parameters for their specific project and data according to their needs, and developers can test their newly developed tools or workflows privately, before adding them as public references. This community-driven effort will increase transparency and reproducibility between MS data analysis workflows, as well as facilitate the development and publication of software workflows in the field.

bioinformatics↗

iDeepLC: chemical structure information yields improved retention time prediction of peptides with unseen modifications

Deep learning has notably advanced the field of liquid chromatography-mass spectrometry-based proteomics. Accurate prediction of peptide retention times significantly enhances our ability to match LC-MS data with the correct peptides and proteins, especially for DIA data. While numerous models predict peptide LC retention times with high accuracy, few can accurately predict the retention times of chemically modified peptides, particularly those with modifications not encountered during model training. In our previously developed DeepLC model, accurate predictions could be made for unseen modifications by leveraging the chemical composition of (modified) residues. Here, however, we present a further enhancement of this model based on chemical structural information. The resulting model, called iDeepLC, shows overall more accurate predictions, and better generalization performance for predicting the retention time of unseen modifications than DeepLC. iDeepLC is freely available as open-source software under the Apache2 license and can be found at https://github.com/CompOmics/iDeepLC.

bioinformatics↗

DeepLC introduces transfer learning for accurate LC retention time prediction and adaptation to substantially different modifications and setups

While LC retention time prediction of peptides and their modifications has proven useful, widespread adoption and optimal performance are hindered by variations in experimental parameters. These variations can render retention time prediction models inaccurate and dramatically reduce the value of predictions for identification, validation, and DIA spectral library generation. To date, mitigation of these issues has been attempted through calibration or by training bespoke models for specific experimental setups, with only partial success. We here demonstrate that transfer learning can successfully overcome these limitations by leveraging pre-trained model parameters. Remarkably, this approach can even fit highly performant models to substantially different peptide modifications and LC conditions than those on which the model was originally trained. This impressive adaptability of transfer learning makes it a highly robust solution for accurate peptide retention time prediction across a very wide variety of imaginable proteomics workflows.

bioinformatics↗

Collisional cross-section prediction for multiconformational peptide ions with IM2Deep

Peptide collisional cross-section (CCS) prediction is complicated by the tendency of peptide ions to exhibit multiple conformations in the gas phase. This adds further complexity to downstream analysis of proteomics data, for example for identification or quantification through feature finding. Here, we present an improved version of IM2Deep that is trained on a carefully curated dataset to predict CCS values of multiconformational peptides. The training data is derived from a large and comprehensive set of publicly available datasets. This comprehensive training dataset together with a tailored architecture allows for the accurate CCS prediction of multiple peptide conformational states. Furthermore, the enhanced IM2Deep model also retains high precision for peptides with a single observed conformation. IM2Deep is publicly available under a permissive open source license at https://github.com/compomics/IM2Deep.

bioinformatics↗