bioRxiv ScienceSearch

Biology subjects

Kertesz-Farkas, A.

Publications and source records attributed to Kertesz-Farkas, A..

2 recordsLinked to original sources

Tailor: non-parametric and rapid score calibration method for database search-based peptide identification in shotgun proteomics

Peptide-spectrum-match (PSM) scores used in database searching are calibrated to spectrum- or spectrum-peptide-specific null distributions. Some calibration methods rely on specific assumptions and use analytical models (e.g. binomial distributions), whereas other methods utilize exact empirical null distributions. The former may be inaccurate because of unjustified assumptions, while the latter are accurate, albeit computationally exhaustive. Here, we introduce a novel, non-parametric, heuristic PSM score calibration method, called Tailor, which calibrates PSM scores by dividing it with the top 100-quantile of the empirical, spectrum-specific null distributions (i.e. the score with an associated p-value of 0.01 at the tail, hence the name) observed during database searching. Tailor does not require any optimization steps or long calculations; it does not rely on any assumptions on the form of the score distribution, it works with any score functions with high- and low-resolution information. In our benchmark, we re-calibrated the match scores of XCorr from Crux, HyperScore scores from X!Tandem, and the p-values from OMSSA with Tailor method, and obtained more spectrum annotation than with raw scores at any false discovery rate level. Moreover, Tailor provided slightly more annotations than E-values of X!Tandem and OMSSA and approached the performance of the computationally exhaustive exact p-value method for XCorr on spectrum datasets containing low-resolution fragmentation information (MS2) around 20-150 times faster. On high-resolution MS2 datasets, the Tailor method with XCorr achieved state-of-the-art performance, and produced more annotations than the well-calibrated Res-ev score around 50-80 times faster.\n\nGraphical TOC Entry\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC=\"FIGDIR/small/831776v1_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (9K):\norg.highwire.dtl.DTLVardef@10e13cdorg.highwire.dtl.DTLVardef@13638e5org.highwire.dtl.DTLVardef@d17e81org.highwire.dtl.DTLVardef@1c86552_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics

Test-time augmentation for deep learning-based cell segmentation on microscopy images

Recent advancements in deep learning have revolutionized the way microscopy images of cells are processed. Deep learning network architectures have a large number of parameters, thus, in order to reach high accuracy, they require massive amount of annotated data. A common way of improving accuracy builds on the artificial increase of the training set by using different augmentation techniques. A less common way relies on test-time augmentation (TTA) which yields transformed versions of the image for prediction and the results are merged. In this paper we describe incorporating the test-time argumentation prediction method into two major segmentation approaches used in the single-cell analysis of microscopy images, namely semantic segmentation using U-Net and instance segmentation using Mask R-CNN models. Our findings show that even using only simple test-time augmentations, such as rotation or flipping and proper merging methods, will result in significant improvement of prediction accuracy. We utilized images of tissue and cell cultures from the Data Science Bowl (DSB) 2018 nuclei segmentation competition and other sources. Additionally, boosting the highest-scoring method of the DSB with TTA, we could further improve and our method has reached an ever-best score at the DSB.

bioinformatics