bioRxiv ScienceSearch

Biology subjects

Parmigiani, G.

Publications and source records attributed to Parmigiani, G..

3 recordsLinked to original sources

The impact of different sources of heterogeneity on loss of accuracy from genomic prediction models

Cross-study validation (CSV) of prediction models is an alternative to traditional cross-validation (CV) in domains where multiple comparable datasets are available. Although many studies have noted potential sources of heterogeneity in genomic studies, to our knowledge none have system atically investigated their intertwined impacts on prediction accuracy across studies. We employ a hybrid parametric/non-parametric bootstrap method to realistically simulate publicly available compendia of microarray, RNA-seq, and whole metagenome shotgun (WMS) microbiome studies of health outcomes. Three types of heterogeneity between studies are manipulated and studied: imbalances in the prevalence of clinical and pathological covariates, 2) differences in gene covariance that could be caused by batch, platform, or tumor purity effects, and 3) differences in the \"true\" model that associates gene expression and clinical factors to outcome. We assess model accuracy while altering these factors. Lower accuracy is seen in CSV than in CV. Surprisingly, heterogeneity in known clinical covariates and differences in gene covariance structure have very limited contributions in the loss of accuracy when validating in new studies. However, forcing identical generative models greatly reduces the within/across study difference. These results, observed consistently for multiple disease outcomes and omics platforms, suggest that the most easily identifiable sources of study heterogeneity are not necessarily the primary ones that undermine the ability to accurately replicate the accuracy of omics prediction models in new studies. Unidentified heterogeneity, such as could arise from unmeasured confounding, may be more important.

bioinformatics

Consensus on Molecular Subtypes of Ovarian Cancer

INTRODUCTIONVarious computational methods for gene expression-based subtyping of high-grade serous (HGS) ovarian cancer have been proposed. This resulted in the identification of molecular subtypes that are based on different datasets and were differentially validated, making it difficult to achieve consensus on which definitions to use in follow-up studies. We assess three major subtype classifiers for their robustness and association to outcome by a meta-analysis of publicly available expression data, and provide a classifier that represents their consensus.\n\nMETHODSWe use a compendium of 15 microarray datasets consisting of 1,774 HGS ovarian tumors to assess 1) concordance between published subtyping algorithms, 2) robustness of those algorithms to re-clustering across datasets, and 3) association of subtypes with overall survival. A consensus classifier is trained on concordantly classified samples, and validated by leave-one-dataset-out validation.\n\nRESULTSEach subtyping classifier identified subsets significantly differing in overall survival, but were not robust to re-fitting in independent datasets and grouped only approximately one third of patients concordantly into four subtypes. We propose a consensus classifier to identify the minority of unambiguously classifiable tumors across multiple gene expression platforms, using a 100-gene signature. The resulting consensus subtypes correlate with patient age, survival, tumor purity, and lymphocyte infiltration.\n\nCONCLUSIONSOur analysis demonstrates that most HGS ovarian cancers are not able to be subtyped. A minority of tumors can be classified and our proposed consensus classifier consolidates and improves on the robustness of three previously proposed subtype classifiers. It provides reliable stratification of patients with HGS ovarian tumors of clearly defined subtype, and will assist in studying the role of polyclonality in the majority of tumors that are not robustly classifiable.

cancer biology

Transcriptome Deconvolution of Heterogeneous Tumor Samples with Immune Infiltration

Transcriptomic deconvolution in cancer and other heterogeneous tissues remains challenging. Available methods lack the ability to estimate both component-specific proportions and expression profiles for individual samples. We present DeMixT, a new tool to deconvolve high dimensional data from mixtures of more than two components. DeMixT implements an iterated conditional mode algorithm and a novel gene-set-based component merging approach to improve accuracy. In a series of experimental validation studies and application to TCGA data, DeMixT showed high accuracy. Improved deconvolution is an important step towards linking tumor transcriptomic data with clinical outcomes. An R package, scripts and data are available: https://github.com/wwylab/DeMixT/.

bioinformatics