bioRxiv ScienceSearch

Biology subjects

Lin, D.-Y.

Publications and source records attributed to Lin, D.-Y..

2 recordsLinked to original sources

MOVIE: Multi-Omics VIsualization of Estimated contributions

SummaryThe growth of multi-omics datasets has given rise to many methods for identifying sources of common variation across data types. The unsupervised nature of these methods makes it difficult to evaluate their performance. We present MOVIE, Multi-Omics Visualization of Estimated contributions, as a framework for evaluating the degree of overfitting and the stability of unsupervised multi-omics methods. MOVIE plots the contributions of one data type against another to produce contribution plots, where contributions are calculated for each subject and each data type from the results of each multi-omics method. The usefulness of MOVIE is demonstrated by applying existing multi-omics methods to permuted null data and breast cancer data from The Cancer Genome Atlas. Contribution plots indicated that principal components-based Canonical Correlation Analysis overfit null data, while Sparse multiple Canonical Correlation Analysis and Multi-Omics Factor Analysis provided stable results with high specificity for both the real and permuted null datasets.\n\nAvailabilityMOVIE is available as an R package at https://github.com/mccabes292/movie\n\nContactmilove@email.unc.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

An Integrative Boosting Approach for Predicting Survival Time With Multiple Genomics Platforms

Recent technological advances have made it possible to collect multiple types of genomics data on the same set of patients. It is of great interest to integrate multiple genomics data types together for predicting disease outcomes. We propose a variable selection method, termed Integrative Boosting (I-Boost), that makes proper use of all available clinical and genomics data in predicting individual patient survival time. Through simulation studies and applications to data sets from The Cancer Genome Atlas, we demonstrate that I-Boost provides substantially higher prediction accuracy than existing variable selection methods. Using I-Boost, we show that (1) the integration of multiple genomics platforms with clinical variables significantly improves the prediction accuracy for survival time over the use of clinical variables alone; (2) gene expression values are typically more prognostic of survival time than other genomics data types; and (3) gene modules/signatures are at least as prognostic as the collection of individual gene expression data.

bioinformatics