bioRxiv ScienceSearch

Biology subjects

Yasin Şenbabaoğlu

Publications and source records attributed to Yasin Şenbabaoğlu.

2 recordsLinked to original sources

A multi-method approach for proteomic network inference in 11 human cancers

Protein expression and post-translational modification levels are tightly regulated in neoplastic cells to maintain cellular processes known as cancer hallmarks. The first Pan-Cancer initiative of The Cancer Genome Atlas (TCGA) Research Network has aggregated protein expression profiles for 3,467 patient samples from 11 tumor types using the antibody based reverse phase protein array (RPPA) technology. The resultant proteomic data can be utilized to computationally infer protein-protein interaction (PPI) networks and to study the commonalities and differences across tumor types. In this study, we compare the performance of 13 established network inference methods in their capacity to retrieve literature-curated pathway interactions from RPPA data. We observe that no single method has the best performance in all tumor types, but a group of six methods, including diverse techniques such as correlation, mutual information, and regression, consistently rank highly among the tested methods. A consensus network from this high-performing group reveals that signal transduction events involving receptor tyrosine kinases (RTKs), the RAS/MAPK pathway, and the PI3K/AKT/mTOR pathway, as well as innate and adaptive immunity signaling, are the most significant PPIs shared across all tumor types. Our results illustrate the utility of the RPPA platform as a tool to study proteomic networks in cancer.\n\nAvailabilityPPI networks from the TCGA or user-provided data can be visualized with the ProtNet web application at http://www.sanderlab.org/protnet/.

Bioinformatics

A reassessment of consensus clustering for class discovery

Consensus clustering (CC) is an unsupervised class discovery method widely used to study sample heterogeneity in high-dimensional datasets. It calculates \"consensus rate\" between any two samples as how frequently they are grouped together in repeated clustering runs under a certain degree of random perturbation. The pairwise consensus rates form a between-sample similarity matrix, which has been used (1) as a visual proof that clusters exist, (2) for comparing stability among clusters, and (3) for estimating the optimal number (K) of clusters. However, the sensitivity and specificity of CC have not been systemically studied. To assess its performance, we investigated the most common implementations of CC; and compared CC with other popular methods that also focus on cluster stability and estimation of K. We evaluated these methods using simulated datasets with either known structure or known absence of structure. Our results showed that (1) CC was able to divide randomly generated unimodal data into pre-specified numbers of clusters, and was able to show apparent stability of these chance partitions of known cluster-less data; (2) for data with known structure, the proportion of ambiguously clustered (PAC) pairs infers the known number of clusters more reliably than several commonly used K estimating methods; and (3) validation of the optimal K by choosing the most discriminant genes from the discovery cohort and applying them in an independent cohort often exaggerates the confidence in K due to inherent gene-gene correlations among the selected genes. While these results do not yet prove that any of the published studies using CC has generated false positive findings, they show that datasets with subtle or no structure are fully capable of producing strong evidence of consensus clustering. We therefore recommend caution is using CC in class discovery and validation.

Bioinformatics