Accurate peptide fragmentation predictions allow data driven approaches to replace and improve upon proteomics search engine scoring functions
The use of post-processing tools to maximize the information gained from a proteomics search engine is widely accepted and used by the community, with the most notable example being Percolator - a semi-supervised machine learning model which learns a new scoring function for a given dataset. The usage of such tools is however bound to the search engines scoring scheme, which doesnt always make full use of the intensity information present in a spectrum. By leveraging another machine learning-based tool, MS2PIP, we aim to overcome this obstacle. MS2PIP predicts fragment ion peak intensities. We show how comparing these intensities to annotated experimental spectra by calculating direct similarity metrics rather than the more common peak counting or explained intensities summing provides enough information for a tool such as Percolator to accurately separate two classes of PSMs, recovering more information out of the data while maintaining control of statistics such as the false discovery rate.