bioRxiv ScienceSearch

Biology subjects

Buyukozkan, M.

Publications and source records attributed to Buyukozkan, M..

2 recordsLinked to original sources

SGI: Automatic clinical subgroup identification in omics datasets

SummaryThe Subgroup Identification (SGI) toolbox provides an algorithm to automatically detect clinical subgroups of samples in large-scale omics datasets. It is based on hierarchical clustering trees in combination with a specifically designed association testing and visualization framework that can process an arbitrary number of clinical parameters and outcomes in a systematic fashion. A multi-block extension allows for the simultaneous use of multiple omics datasets on the same samples. In this paper, we describe the functionality of the toolbox and demonstrate an application example on a blood metabolomics dataset with various clinical biochemistry readouts in a type 2 diabetes case-control study. Availability and implementationSGI is an open-source package implemented in R. Package source codes and hands-on tutorials are available at https://github.com/krumsieklab/sgi. The QMdiab metabolomics data is included in the package and can be downloaded from https://doi.org/10.6084/m9.figshare.5904022.

bioinformatics

PRER: A Patient Representation with Pairwise Relative Expression of Proteins on Biological Networks

Changes in protein and gene expression levels are often used as features to predictive models such as survival prediction. A common strategy to aggregate information on individual proteins is to integrate the expression information with biological networks. We propose a novel patient representation in this work where we integrate proteins expression levels with the protein-protein interaction (PPI) networks. Patient representation with PRER (Pairwise Relative Expressions with Random walks) uses the neighborhood of a protein to capture the dysregulation patterns in protein abundance. Specifically, PRER computes a feature vector for a patient by comparing the source proteins protein expression level with other proteins levels in its neighborhood. This neighborhood of the source protein is derived using a biased random-walk strategy on the network. We test PRERs performance through a survival prediction task in 10 different cancers using random forest survival models. PRER representation yields a statistically significant predictive performance in 9 out of 10 cancer types when compared to a representation based on individual protein expression. We also identify important proteins that are not important in the models trained with the expression values but emerge as predictive in models trained with PRER features. The set of identified relations provides a valuable collection of biomarkers with high prognostic value. PRER representation can be used for other complex diseases and prediction tasks that use molecular expression profiles as input. PRER is freely available at: https://github.com/hikuru/PRER

systems biology