bioRxiv ScienceSearch

Biology subjects

Edmands, W. M. B.

Publications and source records attributed to Edmands, W. M. B..

2 recordsLinked to original sources

Data-adaptive pipeline for filtering and normalizing metabolomics data.

IntroductionUntargeted metabolomics datasets contain large proportions of uninformative features and are affected by a variety of nuisance technical effects that can bias subsequent statistical analyses. Thus, there is a need for versatile and data-adaptive methods for filtering and normalizing data prior to investigating the underlying biological phenomena.\n\nObjectivesHere, we propose and evaluate a data-adaptive pipeline for metabolomics data that are generated by liquid chromatography-mass spectrometry platforms.\n\nMethodsOur data-adaptive pipeline includes novel methods for filtering features based on blank samples, proportions of missing values, and estimated intra-class correlation coefficients. It also incorporates a variant of k-nearest-neighbor imputation of missing values. Finally, we adapted an RNA-Seq approach and R package, scone, to select an appropriate normalization scheme for removing unwanted variation from metabolomics datasets.\n\nResultsUsing two metabolomics datasets that were generated in our laboratory from samples of human blood serum and neonatal blood spots, we compared our data-adaptive pipeline with a traditional filtering and normalization scheme. The data-adaptive approach outperformed the traditional pipeline in almost all metrics related to removal of unwanted variation and maintenance of biologically relevant signatures. The R code for running the data-adaptive pipeline is provided with an example dataset at https://github.com/courtneyschiffman/Data-adaptive-metabolomics.\n\nConclusionOur proposed data-adaptive pipeline is intuitive and effectively reduces technical noise from untargeted metabolomics datasets. It is particularly relevant for interrogation of biological phenomena in data derived from complex matrices associated with biospecimens.

bioinformatics

simExTargId: An R package for real-time LC-MS metabolomic data analysis, instrument failure/drift notification and MS2 target identification

The simExTargId R package provides real-time, autonomous, within-laboratory data analysis during a metabolomic LC-MS1-profiling experiment. Of concern to metabolomic investigators are instrumentation failure (especially for precious samples), outlier identification, instrument signal attenuation and pre-emptive feature identification for MS2 fragmentation.\n\nSimExTargId allows observation of an experiment in progress with PCA plot and peak table outputs and also two shiny applications targetId for MS2 target identification and peakMonitor for signal attenuation monitoring. SimExTargId is ideally utilised on a (temporarily) dedicated workstation or server which is networked to a LC-MS data directory. Features include: email notification for instrument stoppage/drift, file format conversion, peak-picking, pre-processing, PCA-based outlier identification and statistical analysis. Additional MS1/MS2 experiments can be concatenated to a worklist or cleaning/recalibration undertaken if instrument drift is observed.\n\nAll source code and a vignette with example data are available on GitHub https://github.com/WMBEdmands/simExTargId/.\n\nContact: edmandsw@berkeley.edu

bioinformatics