bioRxiv ScienceSearch

Biology subjects

Moseley, H. N. B.

Publications and source records attributed to Moseley, H. N. B..

4 recordsLinked to original sources

Advances in Gene Ontology Utilization Improve Statistical Power of Annotation Enrichment

Gene-annotation enrichment is a common method for utilizing ontology-based annotations in these gene and gene-product centric knowledgebases. Effective utilization of these annotations requires inferring semantic linkages by tracing paths through the ontology through edges in the ontological graph, referred to as relations. However, some relations are semantically problematic with respect to scope, necessitating their omission lest erroneous term mappings occur. To address these issues, we present GOcats, a novel tool that organizes the Gene Ontology (GO) into subgraphs representing user-defined concepts, while ensuring that all appropriate relations are congruent with respect to scoping semantics. Here, we demonstrate the improvements in annotation enrichment by re-interpreting edges that would otherwise be omitted by traditional ancestor path-tracing methods.\n\nWe demonstrate that GOcats unique handling of relations improves enrichment over conventional methods in the analysis of two different gene-expression datasets: a breast cancer microarray dataset and several horse cartilage development RNAseq datasets. With the breast cancer microarray dataset, we observed significant improvement (one-sided binomial test p-value=1.86E-25) in 182 of 217 significantly enriched GO terms identified from the conventional path traversal method when GOcats path traversal was used. We also found new significantly enriched terms using GOcats, whose biological relevancy has been experimentally demonstrated elsewhere. Likewise, on the horse RNAseq datasets, we observed a significant improvement in GO term enrichment when using GOcats path traversal: one-sided binomial test p-values range from 1.32E-03 to 2.58E-44.

systems biology

Inferring metabolite interactomes via molecular structure informed Bayesian graphical model selection with an application to coronary artery disease

IntroductionWhile the generation of reference genomes facilitates the elucidation of gene-phenome associations, reference models of the metabolome that are specific to organism, sample type (e.g. plasma, serum, urine, cell-culture), and state (including disease), remain uncommon. In studying heart disease in humans, a reference model describing the relationships between metabolites in plasma has not been determined but would have great utility as a reference for comparing acute disease states such as myocardial infarction.\n\nMaterials and MethodsWe present a methodology for deriving probabilistic models that describe the partial correlation structure of metabolite distributions (\"interactomes\") from metabolomics data. As determining partial correlation structures requires estimating p*(p-1)/2 parameters for p metabolites, the dimension of the search space for parameter values is immense. Consequently, we have developed a Bayesian methodology for the penalized estimation of model parameters in which the magnitude of penalization is drawn from probability distributions with hyperparameters linked to molecular structure similarity. In our work, structural similarity was determined as the Tanimoto coefficient of algorithmically-generated \"atom colors\" that capture the local structure around each atom within each structure. A Gibbs sampler (a Markov chain Monte Carlo technique) was implemented for simulating the posterior distribution of model parameters. We have made software for implementing this methodology publicly available via the R package BayesianGLasso.\n\nResults / ConclusionsFirst, we demonstrate robust performance of our methodology (sensitivity, specificity, and measures of accuracy) for recovering the true underlying partial correlation structure over simulated datasets (with simulated metabolite abundances and simulated known structural similarity). We then present an interactome model for stable heart disease inferred from non-targeted mass spectrometry data via this methodology. Inspection of the local graph topology about cholate reveals probabilistic interactions with other primary bile acids, secondary bile acids, and many steroid hormones sharing the same precursors.

systems biology

GOcats: A tool for categorizing Gene Ontology into subgraphs of user-defined concepts

Gene Ontology is used extensively in scientific knowledgebases and repositories to organize the wealth of available biological information. However, interpreting annotations derived from differential gene lists is difficult without manually sorting into higher-order categories. To address these issues, we present GOcats, a novel tool that organizes the Gene Ontology (GO) into subgraphs representing user-defined concepts, while ensuring that all appropriate relations are congruent with respect to scoping semantics. We tested GOcats performance using subcellular location categories to mine annotations from GO-utilizing knowledgebases and evaluating their accuracy against immunohistochemistry datasets in the Human Protein Atlas (HPA).\n\nIn comparison to mappings generated from UniProts controlled vocabulary and from GO slims via OWLTools Map2Slim, GOcats outperforms these methods without reliance on a human-curated set of GO terms. By identifying and properly defining relations with respect to semantic scope, GOcats can use traditionally problematic relations without encountering erroneous term mapping. We applied GOcats in the comparison of HPA-sourced knowledgebase annotations to experimentally-derived annotations provided by HPA directly. During the comparison, GOcats improved correspondence between the annotation sources by adjusting semantic granularity. Utilized in this way, GOcats can perform an accurate knowledgebase-level evaluation of curated HPA-based annotations.

systems biology

High Peak Density Artifacts in Fourier Transform Mass Spectra and their Effects on Data Analysis

Fourier-transform mass spectrometry (FT-MS) allows for the high-throughput and high-resolution detection of thousands of metabolites. Observed spectral features (peaks) that are not isotopologues do not directly correspond to known compounds and cannot be placed into existing metabolic networks. Spectral artifacts account for many of these unidentified peaks, and misassignments made to these artifact peaks can create large interpretative errors. Without accurate identification of artifactual features and correct assignment of real features, discerning their roles within living systems is effectively impossible.\n\nWe have observed three types of artifacts unique to FT-MS that often result in regions of abnormally high peak density (HPD), which we collectively refer to as HPD artifacts: i) fuzzy sites representing small regions of m/z space with a fuzzy appearance due to the extremely high number of peaks present; ii) ringing due to a very intense peak producing side bands of decreasing intensity that are symmetrically distributed around the main peak; and iii) partial ringing where only a subset of the side bands are observed for an intense peak. Fuzzy sites and partial ringing appear to be novel artifacts previously unreported in the literature and we hypothesize that all three artifact types derive from Fourier transformation-based issues. In some spectra, these artifacts account for roughly a third of the peaks present in the given spectrum. We have developed a set of tools to detect these artifacts and approaches to mitigate their effects on downstream analyses.

systems biology