bioRxiv ScienceSearch

Biology subjects

Gama-Castro, S.

Publications and source records attributed to Gama-Castro, S..

3 recordsLinked to original sources

Towards a unified resource for transcriptional regulation in Escherichia coli K-12: Incorporating high throughput-generated binding data within the classic framework of regulation of initiation of transcription in RegulonDB.

Our understanding of the regulation of gene expression has been strongly benefited by the availability of high throughput technologies that enable questioning the whole genome for the binding of specific transcription factors and expression profiles. In the case of genome models, such as Escherichia coli K-12, this knowledge needs to be integrated with the legacy of accumulated genetics and molecular biology pre-genomic knowledge in order to attain deeper levels in the understanding of their biology. In spite of the several repositories and curated databases, there is no effort, nor electronic site yet, to comprehensively integrate the available knowledge from all these different sources around the regulation of gene expression of E. coli K-12. In this paper, we describe a first effort to expand RegulonDB, the database containing the rich legacy of decades of classic molecular biology experiments supporting what we know about gene regulation and operon organization in E. coli K-12, to include the genome-wide data set collections from 25 ChIP and 18 gSELEX publications, respectively, in addition to around 60 expression profiles used in their curation. Three essential features for the integration of this information coming from different methodological approaches are; first, a controlled vocabulary within an ontology for precisely defining growth conditions, second, the criteria to separate elements with enough evidence to consider them involved in gene regulation from isolated sites, and third, an expanded computational model supporting this knowledge. Altogether, this constitutes the basis for adequately gathering and enabling the comparisons and integration strongly needed to manage and access such wealth of knowledge. This version of RegulonBD is a first step toward what should become the unifying access point for current and future knowledge on gene regulation in E. coli K-12. Furthermore, this model platform and associated methodologies and criteria, can well be emulated for gathering knowledge on other microbial organisms.

systems biology

Similarity corpus on microbial transcriptional regulation

The ability to express the same meaning in different ways is a well known property of natural language. This amazing property is the source of major difficulties in natural language processing. Given the constant increase in published literature, its curation and information extraction would strongly benefit by efficient automatic processes, for which, corpora of sentences evaluated by experts is a valuable resource. Given our interest in applying such approaches to the benefit of curation of the biomedical literature, specifically about gene regulation in microbial organisms, we decided to build a corpus with graded textual similarity evaluated by curators, and designed specifically oriented to our purposes. Based on the predefined statistical power of future analyses, we defined features of the design including sampling, selection criteria, balance, and size among others. A non-fully crossed-design was performed for each pair of sentences by 3 evaluators from 7 different groups, adapting the SEMEVAL scale to our goals in four successive iterative sessions with a clear improvement in the consensuated guidelines and inter-rater-reliability results. Alternatives for the corpus evaluation are widely discussed. To the best of our knowledge this is the first similarity corpus in this domain of knowledge. We have initiated its incorporation in our research towards high throughput curation strategies based in natural language processing.

bioinformatics

MCO: towards an ontology and unified vocabulary for a framework-based annotation of microbial growth conditions

MotivationA major component in our understanding of the biology of an organism is the mapping of its genotypic potential into the repertoire of its phenotypic expression profiles. This genotypic to phenotypic mapping is executed by the machinery of gene regulation that turns genes on and off, which in microorganisms is essentially studied by changes in growth conditions and genetic modifications. Although many efforts have been made to systematize the annotation of experimental conditions in microbiology, the available annotation is not based on a consistent and controlled vocabulary for the unambiguous description of growth conditions, making difficult the identification of biologically meaningful comparisons of knowledge generated in different experiments or laboratories, a task urgently needed given the massive amounts of data generated by high throughput (HT) technologies.\n\nResultsWe curated terms related to experimental conditions that affect gene expression in E. coli K-12. Since this is the best studied microorganism, the collected terms are the seed for the first version of the Microbial Conditions Ontology (MCO), a controlled and structured vocabulary that can be expanded to annotate microbial conditions in general. Moreover, we developed an annotation framework using the MCO terms to describe experimental conditions, providing the foundation to identify regulatory networks that operate under a particular condition. MCO supports comparisons of HT-derived data from different repositories. In this sense, we started to map common RegulonDB terms and Colombos bacterial expression compendia terms to MCO.\n\nAvailability and ImplementationAs far as we know, MCO is the first ontology for growth conditions of any bacterial organism and it is available at http://regulondb.ccg.unam.mx/. Furthermore, we will disseminate MCO throughout the Open Biomedical Ontology (OBO) Foundry in order to set a standard for the annotation of gene expression data derived from conventional as well as HT experiments in E. coli and other microbial organisms. This will enable the comparison of data from diverse data sources.\n\nContactsgama@ccg.unam.mx, collado@ccg.unam.mx

bioinformatics