bioRxiv ScienceSearch

Biology subjects

Mejia-Almonte, C.

Publications and source records attributed to Mejia-Almonte, C..

2 recordsLinked to original sources

Similarity corpus on microbial transcriptional regulation

The ability to express the same meaning in different ways is a well known property of natural language. This amazing property is the source of major difficulties in natural language processing. Given the constant increase in published literature, its curation and information extraction would strongly benefit by efficient automatic processes, for which, corpora of sentences evaluated by experts is a valuable resource. Given our interest in applying such approaches to the benefit of curation of the biomedical literature, specifically about gene regulation in microbial organisms, we decided to build a corpus with graded textual similarity evaluated by curators, and designed specifically oriented to our purposes. Based on the predefined statistical power of future analyses, we defined features of the design including sampling, selection criteria, balance, and size among others. A non-fully crossed-design was performed for each pair of sentences by 3 evaluators from 7 different groups, adapting the SEMEVAL scale to our goals in four successive iterative sessions with a clear improvement in the consensuated guidelines and inter-rater-reliability results. Alternatives for the corpus evaluation are widely discussed. To the best of our knowledge this is the first similarity corpus in this domain of knowledge. We have initiated its incorporation in our research towards high throughput curation strategies based in natural language processing.

bioinformatics

MCO: towards an ontology and unified vocabulary for a framework-based annotation of microbial growth conditions

MotivationA major component in our understanding of the biology of an organism is the mapping of its genotypic potential into the repertoire of its phenotypic expression profiles. This genotypic to phenotypic mapping is executed by the machinery of gene regulation that turns genes on and off, which in microorganisms is essentially studied by changes in growth conditions and genetic modifications. Although many efforts have been made to systematize the annotation of experimental conditions in microbiology, the available annotation is not based on a consistent and controlled vocabulary for the unambiguous description of growth conditions, making difficult the identification of biologically meaningful comparisons of knowledge generated in different experiments or laboratories, a task urgently needed given the massive amounts of data generated by high throughput (HT) technologies.\n\nResultsWe curated terms related to experimental conditions that affect gene expression in E. coli K-12. Since this is the best studied microorganism, the collected terms are the seed for the first version of the Microbial Conditions Ontology (MCO), a controlled and structured vocabulary that can be expanded to annotate microbial conditions in general. Moreover, we developed an annotation framework using the MCO terms to describe experimental conditions, providing the foundation to identify regulatory networks that operate under a particular condition. MCO supports comparisons of HT-derived data from different repositories. In this sense, we started to map common RegulonDB terms and Colombos bacterial expression compendia terms to MCO.\n\nAvailability and ImplementationAs far as we know, MCO is the first ontology for growth conditions of any bacterial organism and it is available at http://regulondb.ccg.unam.mx/. Furthermore, we will disseminate MCO throughout the Open Biomedical Ontology (OBO) Foundry in order to set a standard for the annotation of gene expression data derived from conventional as well as HT experiments in E. coli and other microbial organisms. This will enable the comparison of data from diverse data sources.\n\nContactsgama@ccg.unam.mx, collado@ccg.unam.mx

bioinformatics