bioRxiv ScienceSearch

Biology subjects

Do, K.

Publications and source records attributed to Do, K..

2 recordsLinked to original sources

Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies

BACKGROUNDUntargeted mass spectrometry (MS)-based metabolomics data often contain missing values that reduce statistical power and can introduce bias in epidemiological studies. However, a systematic assessment of the various sources of missing values and strategies to handle these data has received little attention. Missing data can occur systematically, e.g. from run day-dependent effects due to limits of detection (LOD); or it can be random as, for instance, a consequence of sample preparation.\n\nMETHODSWe investigated patterns of missing data in an MS-based metabolomics experiment of serum samples from the German KORA F4 cohort (n = 1750). We then evaluated 31 imputation methods in a simulation framework and biologically validated the results by applying all imputation approaches to real metabolomics data. We examined the ability of each method to reconstruct biochemical pathways from data-driven correlation networks, and the ability of the method to increase statistical power while preserving the strength of established genetically metabolic quantitative trait loci.\n\nRESULTSRun day-dependent LOD-based missing data accounts for most missing values in the metabolomics dataset. Although multiple imputation by chained equations (MICE) performed well in many scenarios, it is computationally and statistically challenging. K-nearest neighbors (KNN) imputation on observations with variable pre-selection showed robust performance across all evaluation schemes and is computationally more tractable.\n\nCONCLUSIONMissing data in untargeted MS-based metabolomics data occur for various reasons. Based on our results, we recommend that KNN-based imputation is performed on observations with variable pre-selection since it showed robust results in all evaluation schemes.\n\nKey messagesO_LIUntargeted MS-based metabolomics data show missing values due to both batch-specific LOD-based and non-LOD-based effects.\nC_LIO_LIStatistical evaluation of multiple imputation methods was conducted on both simulated and real datasets.\nC_LIO_LIBiological evaluation on real data assessed the ability of imputation methods to preserve statistical inference of biochemical pathways and correctly estimate effects of genetic variants on metabolite levels.\nC_LIO_LIKNN-based imputation on observations with variable pre-selection and K = 10 showed robust performance for all data scenarios across all evaluation schemes.\nC_LI

systems biology

MatchMiner: An open source computational platform for real-time matching of cancer patients to precision medicine clinical trials using genomic and clinical criteria

BackgroundMolecular profiling of cancers is now routine at many cancer centers, and the number of precision cancer medicine clinical trials, which are informed by profiling, is steadily rising. Additionally, these trials are becoming increasingly complex, often having multiple arms and many genomic eligibility criteria. Currently, it is a challenging for physicians to match patients to relevant clinical trials using the patients genomic profile, which can lead to missed opportunities. Automated matching against uniformly structured and encoded genomic eligibility criteria is essential to keep pace with the complex landscape of precision medicine clinical trials.\n\nResultsTo meet these needs, we built and deployed an automated clinical trial matching platform called MatchMiner at the Dana-Farber Cancer Institute (DFCI). The platform has been integrated with Profile, DFCIs enterprise genomic profiling project, which contains tumor profile data for >20,000 patients, and has been made available to physicians across the Institute. As no current standard exists for encoding clinical trial eligibility criteria, a new language called Clinical Trial Markup Language (CTML) was developed, and over 178 genomically-driven clinical trials were encoded using this language. The platform is open source and freely available for adoption by other institutions.\n\nConclusionMatchMiner is the first open platform developed to enable computational matching of patient-specific genomic profiles to precision cancer medicine clinical trials. Creating MatchMiner required developing clinical trial eligibility standards to support genome-driven matching and developing intuitive interfaces to support practical use-cases. Given the complexity of tumor profiling and the rapidly changing multi-site nature of genome-driven clinical trials, open source software is the most efficient, scalable, and economical option for matching cancer patients to clinical trials.

bioinformatics