bioRxiv ScienceSearch

Biology subjects

Waldron, L.

Publications and source records attributed to Waldron, L..

5 recordsLinked to original sources

The impact of different sources of heterogeneity on loss of accuracy from genomic prediction models

Cross-study validation (CSV) of prediction models is an alternative to traditional cross-validation (CV) in domains where multiple comparable datasets are available. Although many studies have noted potential sources of heterogeneity in genomic studies, to our knowledge none have system atically investigated their intertwined impacts on prediction accuracy across studies. We employ a hybrid parametric/non-parametric bootstrap method to realistically simulate publicly available compendia of microarray, RNA-seq, and whole metagenome shotgun (WMS) microbiome studies of health outcomes. Three types of heterogeneity between studies are manipulated and studied: imbalances in the prevalence of clinical and pathological covariates, 2) differences in gene covariance that could be caused by batch, platform, or tumor purity effects, and 3) differences in the \"true\" model that associates gene expression and clinical factors to outcome. We assess model accuracy while altering these factors. Lower accuracy is seen in CSV than in CV. Surprisingly, heterogeneity in known clinical covariates and differences in gene covariance structure have very limited contributions in the loss of accuracy when validating in new studies. However, forcing identical generative models greatly reduces the within/across study difference. These results, observed consistently for multiple disease outcomes and omics platforms, suggest that the most easily identifiable sources of study heterogeneity are not necessarily the primary ones that undermine the ability to accurately replicate the accuracy of omics prediction models in new studies. Unidentified heterogeneity, such as could arise from unmeasured confounding, may be more important.

bioinformatics

HMP16SData: Efficient Access to the Human Microbiome Project through Bioconductor

Phase 1 of the NIH Human Microbiome Project (HMP) investigated 18 body subsites of 239 healthy American adults, to produce the first comprehensive reference for the composition and variation of the \"healthy\" human microbiome. Publicly-available data sets from amplicon sequencing of two 16S rRNA variable regions, with extensive controlled-access participant data, provide a reference for ongoing microbiome studies. However, utilization of these data sets can be hindered by the complex bioinformatic steps required to access, import, decrypt, and merge the various components in formats suitable for ecological and statistical analysis. The HMP16SData package provides count data for both 16S variable regions, integrated with phylogeny, taxonomy, public participant data, and controlled participant data for authorized researchers, using standard integrative Bioconductor data objects. By removing bioinformatic hurdles of data access and management, HMP16SData enables epidemiologists with only basic R skills to quickly analyze HMP data.

bioinformatics

Consensus on Molecular Subtypes of Ovarian Cancer

INTRODUCTIONVarious computational methods for gene expression-based subtyping of high-grade serous (HGS) ovarian cancer have been proposed. This resulted in the identification of molecular subtypes that are based on different datasets and were differentially validated, making it difficult to achieve consensus on which definitions to use in follow-up studies. We assess three major subtype classifiers for their robustness and association to outcome by a meta-analysis of publicly available expression data, and provide a classifier that represents their consensus.\n\nMETHODSWe use a compendium of 15 microarray datasets consisting of 1,774 HGS ovarian tumors to assess 1) concordance between published subtyping algorithms, 2) robustness of those algorithms to re-clustering across datasets, and 3) association of subtypes with overall survival. A consensus classifier is trained on concordantly classified samples, and validated by leave-one-dataset-out validation.\n\nRESULTSEach subtyping classifier identified subsets significantly differing in overall survival, but were not robust to re-fitting in independent datasets and grouped only approximately one third of patients concordantly into four subtypes. We propose a consensus classifier to identify the minority of unambiguously classifiable tumors across multiple gene expression platforms, using a 100-gene signature. The resulting consensus subtypes correlate with patient age, survival, tumor purity, and lymphocyte infiltration.\n\nCONCLUSIONSOur analysis demonstrates that most HGS ovarian cancers are not able to be subtyped. A minority of tumors can be classified and our proposed consensus classifier consolidates and improves on the robustness of three previously proposed subtype classifiers. It provides reliable stratification of patients with HGS ovarian tumors of clearly defined subtype, and will assist in studying the role of polyclonality in the majority of tumors that are not robustly classifiable.

cancer biology

Software For The Integration Of Multi-Omics Experiments In Bioconductor

Multi-omics experiments are increasingly commonplace in biomedical research, and add layers of complexity to experimental design, data integration, and analysis. R and Bioconductor provide a generic framework for statistical analysis and visualization, as well as specialized data classes for a variety of high-throughput data types, but methods are lacking for integrative analysis of multi-omics experiments. The MultiAssayExperiment software package, implemented in R and leveraging Bioconductor software and design principles, provides for the coordinated representation of, storage of, and operation on multiple diverse genomics data. We provide all of the multiple omics data for each cancer tissue in The Cancer Genome Atlas (TCGA) as ready-to-analyze MultiAssayExperiment objects, and demonstrate in these and other datasets how the software simplifies data representation, statistical analysis, and visualization. The MultiAssayExperiment Bioconductor package reduces major obstacles to efficient, scalable and reproducible statistical analysis of multi-omics data and enhances data science applications of multiple omics datasets.

bioinformatics

Accessible, curated metagenomic data through ExperimentHub

We present curatedMetagenomicData, a Bioconductor and command-line interface to thousands of metagenomic profiles from the Human Microbiome Project and other publicly available datasets, and ExperimentHub, a platform for convenient cloud-based distribution of data to the R desktop. The resource provides standardized per-participant metadata linked to bacterial, fungal, archaeal, and viral taxonomic abundances, as well as quantitative metabolic functional profiles. The datasets can be immediately analyzed in R or other software with a minimum of bioinformatic expertise and no preprocessing of data. We demonstrate identification of taxonomic/functional correlations, an investigation of gut \"enterotypes\", and a comparison of the accuracy of disease classification from different data types. These documented analyses can be reproduced efficiently on a laptop, without the barriers of working with large-scale, raw sequencing data. The building and expansion of curatedMetagenomicData is based entirely on open source software and pipelines, to facilitate the addition of new microbiome datasets and methods.

bioinformatics