bioRxiv ScienceSearch

Biology subjects

Limsoon Wong

Publications and source records attributed to Limsoon Wong.

2 recordsLinked to original sources

Inverting proteomics analysis provides powerful insight into the peptide/protein conundrum

In proteomics, a large proportion of mass spectrometry (MS) data is ignored due to the lack of, or insufficient statistical evidence for mappable peptides. In reality, only a small fraction of features are expected to be differentially relevant anyway. Mapping spectra to peptides and subsequently, proteins, produces uncertainty at several levels. We propose it is better to analyze proteomic profiling data directly at MS level, and then relate these features to peptides/proteins. In a renal cancer data comprising 12 normal and 12 cancer subjects, we demonstrate that a simple rule-based binning approach can give rise to informative features. We note that the peptides associated with significant spectral bins gave rise to better class separation than the corresponding proteins, suggesting a loss of signal in the peptide-to-protein transition. Additionally, the binning approach sharpens focus on relevant protein splice forms rather than just canonical sequences. Taken together, the inverted raw spectra analysis paradigm, which is realised by the MZ-Bin method described in this article, provides new possibilities and insights, in how MS-data can be interpreted.

Bioinformatics

Overcoming analytical reliability issues in clinical proteomics using rank-based network approaches

Proteomics is poised to play critical roles in clinical research. However, due to limited coverage and high noise, integration with powerful analysis algorithms is necessary. In particular, network-based algorithms can improve selection of reproducible features in spite of incomplete proteome coverage, technical inconsistency or high inter-sample variability. We define analytical reliability on three benchmarks --- precision/recall rates, feature-selection stability and cross-validation accuracy. Using these, we demonstrate the insufficiencies of commonly used Students t-test and Hypergeometric enrichment. Given advances in sample sizes, quantitation accuracy and coverage, we are now able to introduce and evaluate Ranked-Based Network Approaches (RBNAs) for the first time in proteomics. These include SNET (SubNETwork), FSNET (FuzzySNET), PFSNET (PairedFSNET). We also introduce for the first time, PPFSNET(samplePairedPFSNET), which is a paired-sample variant of PFSNET. RBNAs (particularly PFSNET and PPFSNET) excelled on all three benchmarks and can make consistent and reproducible predictions even in the small-sample size scenario (n=4). Given these qualities, RBNAs represent an important advancement in network biology, and is expected to see practical usage, particularly in clinical biomarker and drug target prediction.

Bioinformatics