bioRxiv Science⌕ Search

Biology subjects

Karp, P. D.

Publications and source records attributed to Karp, P. D..

5 recordsLinked to original sources

Improved BioCyc Operon Prediction: Revisiting theOperon Prediction Problem

IntroductionOperon prediction is a valuable component of microbial-genome annotation because operon organization can yield inferences about gene function, and because knowledge of operon structure can aid the interpretation of gene expression data. MethodsWe present a number of improvements to the existing Pathway Tools operon predictor based mostly on 7 new features that we hypothesized would increase its performance. The new features include shared Gene Ontology biological process terms, similarity of codon usage and GC content, correlated gene expression, and shared protein complex. ResultsWe evaluated the proposed 7 new features and found that the addition of 6 of them improved the performance of the operon predictor from 79.55% to 83.49%, a decrease in error rate of 19.3%. When gene expression data was not included, the accuracy decreased to 82.547, still an improvement of 14.7%. One of the proposed features as well as a previously used feature had no effect and were removed. DiscussionAlthough some of the new features had strong predictive value individually, when combined with the other features they did not have a large impact on predictive accuracy, suggesting that they were not independent from the other features.

genomics↗

The Comparative Genome Dashboard

The Comparative Genome Dashboard is a web-based software tool for interactive exploration of the similarities and differences in gene functions between organisms. It provides a high-level graphical survey of cellular functions, and enables the user to drill down to examine subsystems of interest in greater detail. At its highest level the Comparative Dashboard contains panels for cellular systems such as biosynthesis, energy metabolism, transport, and response to stimulus. Each panel contains a set of bar graphs that plot the numbers of compounds or gene products for each organism across a set of subsystems of that panel. Users can interactively drill down to focus on subsystems of interest and see grids of compounds produced or consumed by each organism, specific GO term assignments, pathway diagrams, and links to more detailed comparison pages. For example, the dashboard enables users to compare the cofactors that a set of organisms can synthesize, the metal ions that they are able to transport, their DNA damage repair capabilities, their biofilm-formation genes, and their viral response proteins. The dashboard enables users to quickly perform comprehensive comparisons at varying levels of detail.

bioinformatics↗

Visual Analysis of Multi-Omics Data

We present a tool for multi-omics data analysis that enables simultaneous visualization of up to four types of omics data on organism-scale metabolic network diagrams. The tools interactive web-based metabolic charts depict the metabolic reactions, pathways, and metabolites of a single organism as described in a metabolic pathway database for that organism; the charts are constructed using automated graphical layout algorithms. The multi-omics visualization facility paints each individual omics dataset onto a different "visual channel" of the metabolic-network diagram. For example, a transcriptomics dataset might be displayed by coloring the reaction arrows within the metabolic chart, while a companion proteomics dataset is displayed as reaction arrow thicknesses, and a complementary metabolomics dataset is displayed as metabolite node colors. Once the network diagrams are painted with omics data, semantic zooming provides more details within the diagram as the user zooms in. Datasets containing multiple time points can be displayed in an animated fashion. The tool will also graph data values for individual reactions or metabolites designated by the user. The user can interactively adjust the mapping from data value ranges to the displayed colors and thicknesses to provide more informative diagrams.

bioinformatics↗

The Genome Explorer Genome Browser

Are two adjacent genes in the same operon? What is the order and spacing between several transcription-factor binding sites? Genome browsers are software data-visualization and exploration tools that enable biologists to answer questions such as these. In this paper we report on a major update to our browser, Genome Explorer, that provides nearly instantaneous scaling and traversing of a genome, enabling users to quickly and easily zoom into an area of interest. The user can rapidly move between scales that depict the entire genome, individual genes, and the sequence; Genome Explorer presents the most relevant detail and context for each scale. By downloading the data for the entire genome to the users web browser and dynamically generating visualizations locally, we enable fine control of zoom and pan functions and real-time redrawing of the visualization, resulting in smoother and more intuitive exploration of a genome than is possible with other browsers. Further, genome features are presented together, in-line, using familiar graphical depictions. In contrast, many other browsers depict genome features using data tracks, which have low information density and can visually obscure the relative positions of features. Genome Explorer diagrams have high information density that provides larger amounts of genome context and sequence information to be presented in a given sized monitor than for tracks-based browsers. Genome Explorer provides optional data tracks for analysis of large-scale datasets and a unique comparative mode that aligns genomes at orthologous genes with synchronized zooming.

bioinformatics↗

Collaborative metabolic curation of an emerging model marine bacterium, Alteromonas macleodii ATCC 27126

Inferring the metabolic capabilities of an organism from its genome is a challenging process, relying on computationally-derived or manually curated metabolic networks. Manual curation can correct mistakes in the draft network and add missing reactions based on the literature, but requires significant expertise and is often the bottleneck for high-quality metabolic reconstructions. Here, we present a synopsis of a community curation workshop for the emerging model marine bacterium Alteromonas macleodii ATCC 27126 and its genome database in BioCyc, focusing on pathways for utilizing organic carbon and nitrogen sources. Due to the scarcity of biochemical information or gene knock-outs, the curation process relied primarily on published growth phenotypes and bioinformatic analyses, including comparisons with related Alteromonas strains. We report full pathways for the utilization of the algal polysaccharides alginate and pectin in contrast to inconclusive evidence for one carbon metabolism and mixed acid fermentation, in accordance with the lack of growth on methanol and formate. Pathways for amino acid degradation are ubiquitous across Alteromonas macleodii strains, yet enzymes in the pathways for the degradation of threonine, tryptophan and tyrosine were not identified. Nucleotide degradation pathways are also partial in ATCC 27126. We postulate that demonstrated growth on nitrate as sole N source proceeds via a nitrate reductase pathway that is a hybrid of known pathways. Our evidence highlights the value of joint and interactive curation efforts, but also shows major knowledge gaps regarding Alteromonas metabolism. The manually-curated metabolic reconstruction is available as a "Tier-2" database on BioCyc. ImportanceMetabolic reconstructions are vital for the systemic understanding of an organisms ecology. Here, we report the outcome of a collaborative, interactive curation workshop to build a curated "metabolic encyclopedia" for Alteromonas macleodii ATCC 27126, a marine heterotrophic bacterium with widespread occurrence. Curating pathways for polysaccharide degradation, one-carbon metabolism, and others closed major knowledge gaps, and identified further avenues of research. Our study highlights how the combination of bioinformatic, genomic and physiological evidence can be harvested into a detailed metabolic model, but also identifies challenges if little experimental data is available for support. Overall, we show how an interactive get-together by a diverse group of scientists can advance the ecological understanding of emerging model bacteria, with relevance for the entire scientific community.

microbiology↗