bioRxiv ScienceSearch

Biology subjects

Trainor, P. J.

Publications and source records attributed to Trainor, P. J..

3 recordsLinked to original sources

Inferring metabolite interactomes via molecular structure informed Bayesian graphical model selection with an application to coronary artery disease

IntroductionWhile the generation of reference genomes facilitates the elucidation of gene-phenome associations, reference models of the metabolome that are specific to organism, sample type (e.g. plasma, serum, urine, cell-culture), and state (including disease), remain uncommon. In studying heart disease in humans, a reference model describing the relationships between metabolites in plasma has not been determined but would have great utility as a reference for comparing acute disease states such as myocardial infarction.\n\nMaterials and MethodsWe present a methodology for deriving probabilistic models that describe the partial correlation structure of metabolite distributions (\"interactomes\") from metabolomics data. As determining partial correlation structures requires estimating p*(p-1)/2 parameters for p metabolites, the dimension of the search space for parameter values is immense. Consequently, we have developed a Bayesian methodology for the penalized estimation of model parameters in which the magnitude of penalization is drawn from probability distributions with hyperparameters linked to molecular structure similarity. In our work, structural similarity was determined as the Tanimoto coefficient of algorithmically-generated \"atom colors\" that capture the local structure around each atom within each structure. A Gibbs sampler (a Markov chain Monte Carlo technique) was implemented for simulating the posterior distribution of model parameters. We have made software for implementing this methodology publicly available via the R package BayesianGLasso.\n\nResults / ConclusionsFirst, we demonstrate robust performance of our methodology (sensitivity, specificity, and measures of accuracy) for recovering the true underlying partial correlation structure over simulated datasets (with simulated metabolite abundances and simulated known structural similarity). We then present an interactome model for stable heart disease inferred from non-targeted mass spectrometry data via this methodology. Inspection of the local graph topology about cholate reveals probabilistic interactions with other primary bile acids, secondary bile acids, and many steroid hormones sharing the same precursors.

systems biology

Wisdom of artificial crowds feature selection in untargeted metabolomics: An application to the development of a blood-based diagnostic test for thrombotic myocardial infarction

IntroductionHeart disease remains a leading cause of global mortality. While acute myocardial infarction (colloquially: heart attack), has multiple proximate causes, proximate etiology cannot be determined by a blood-based diagnostic test. We enrolled a suitable patient cohort and conducted an untargeted quantification of plasma metabolites by mass spectrometry for developing a test that can differentiate between thrombotic MI, non-thrombotic MI, and stable disease. A significant challenge in developing such a diagnostic test is solving the NP-hard problem of feature selection for constructing an optimal statistical classifier.\n\nObjectiveWe employed a Wisdom of Artificial Crowds (WoAC) strategy for solving the feature selection problem and evaluated the accuracy and parsimony of downstream classifiers in comparison with embedded feature selection via the Lasso and Elastic Net.\n\nMaterials and MethodsArtificial Crowd Wisdom was generated via aggregation of the best solutions from independent and diverse genetic algorithm populations that were initialized with bootstrapping and a random subspaces constraint.\n\nResults / ConclusionsWoAC feature selection performed favorably compared to Lasso and Elastic Net solutions. The classifier constructed following WoAC feature selection had a cross-validation estimated misclassification rate of 2.6% as compared to 26.3% via the Lasso and 18.5% via an Elastic Net. The classifier warrants further evaluation as a diagnostic test in an independent cohort.

bioinformatics

Evaluation Of Classifier Performance For Multiclass Phenotype Discrimination In Untargeted Metabolomics

Statistical classification is a critical component of utilizing metabolomics data for examining the molecular determinants of phenotypes and for furnishing diagnostic and prognostic phenotype predictions in medicine. Despite this, a comprehensive and rigorous evaluation of classification techniques for phenotype discrimination given metabolomics data has not been conducted. We conducted such an evaluation using both simulated and real metabolomics data, comparing Partial Least Squares-Discriminant Analysis (PLS-DA), Sparse PLS-DA, Random Forests, Support Vector Machines, and Neural Network classification techniques for discriminating phenotype. We evaluated the techniques on simulated data generated to mimic global untargeted metabolomics data by incorporating realistic block-wise correlation and partial correlation structures for mimicking the correlations and metabolite clustering generated by biological processes. Over the simulation studies, covariance structures, means, and effect sizes were randomly simulated to provide consistent estimates of classifier performance over a wide range of possible scenarios. The presence of non-normal error distributions and the effect of prior-significance filtering (dimension reduction) were evaluated. In each simulation, classifier parameters (such as the number of hidden nodes in a neural network) were tuned by cross-validation to minimize the probability of detecting spurious results due to poorly tuned classifiers. Classifier performance was then evaluated using real clinical metabolomics datasets of varying sample medium, sample size, and experimental design. We report that in the scenarios without a significant presence of non-normal error distributions over metabolite clusters, Neural Network and PLS-DA classifiers performed poorly relative to Sparse PLS-DA (sPLS-DA), Support Vector Machine (SVM), and Random Forest classifiers. When non-normal error distributions were introduced, the performance of PLS-DA classifiers deteriorated further relative to the remaining techniques. Simultaneously, while the relative performance of Neural Network classifiers improved relative to PLS-DA classifiers, Neural Network classifier performance remained poor compared sPLS-DA, SVM, and Random Forest classifiers. Over the real datasets, a trend of better performance of SVM and Random Forest classifier performance was observed.

bioinformatics