bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Subharmonics And Chaos In Simple Periodically-Forced Biomolecular Models

This paper uncovers a remarkable behavior in two biochemical systems that commonly appear as components of signal transduction pathways in systems biology. These systems have globally attracting steady states when unforced, so they might have been considered \"uninteresting\" from a dynamical standpoint. However, when subject to a periodic excitation, strange attractors arise via a period-doubling cascade. Quantitative analyses of the corresponding discrete chaotic trajectories are conducted numerically by computing largest Lyapunov exponents, power spectra, and autocorrelation functions. To gain insight into the geometry of the strange attractors, the phase portraits of the corresponding iterated maps are interpreted as scatter plots for which marginal distributions are additionally evaluated. The lack of entrainment to external oscillations, in even the simplest biochemical networks, represents a level of additional complexity in molecular biology, which has previously been insufficiently recognized but is plausibly biologically important.

systems biology

MORPHOMETRIC IDENTIFICATION OF STEM BORERS Diatraea saccharalis AND Diatraea busckella (Lepidoptera: Crambidae) IN SUGARCANE CROPS (Saccharum officinarum) IN CALDAS DEPARTMENT, COLOMBIA

The sugarcane (Saccharum sp.), of great importance for being one of the most traditional rural agroindustries in Latin America and the Caribbean, as part of the agricultural systems, is vulnerable to increases or reductions in the incidence of pests associated with extreme events of climate change, such as prolonged droughts, hurricanes, heavy and out of season rains, among others, contributing to the increase losses in agricultural production, which forces farmers to make excessive expenditures on pesticides that generally fail to solve the issue. (Vazquez, 2011). The main pest belongs to the Diatraea complex (Vargas et al., 2013; Gallego et al., 1996), a larval stage perforator habit. Different field evaluations have revealed the presence of a species that had not been reported in sugarcane crops, Diatraea busckella, and to corroborate the finding, a method of identification was needed whose advantage was to be quick and also low cost, in this sense, geometric morphometry is a mathematical tool with biological basis (Bookstein, 1991), which allows to decompose the variation resulting from the physiology of individuals of the most stable individuals of the population, product of the genetic component. CLIC (Collecting Landmarks for Identification and Characterization) was used for identification, with reference to the previous right wing (De La Riva et al., 2001; Belen et al., 2004; Schachter-Broide et al., 2004; Dvorak et al., 2006; Soto Vivas et al., 2007). Wing morphometry was performed using generalized Procrustes analysis (Rohlf and Marcus, 1993). The analysis clearly differentiated between D. busckella and D. saccharalis, eliminating the environmental factors that could generate some level of error, being considered a support tool that validates the molecular biology processes for the identification of organisms.

ecology

Wikidata as a semantic framework for the Gene Wiki initiative

Open biological data is distributed over many resources making it challenging to integrate, to update and to disseminate quickly. Wikidata is a growing, open community database which can serve this purpose and also provides tight integration with Wikipedia.\n\nIn order to improve the state of biological data, facilitate data management and dissemination, we imported all human and mouse genes, and all human and mouse proteins into Wikidata. In total, 59,530 human genes and 73,130 mouse genes have been imported from NCBI and 27,662 human proteins and 16,728 mouse proteins have been imported from the Swissprot subset of UniProt. As Wikidata is open and can be edited by anybody, our corpus of imported data serves as the starting point for integration of further data by scientists, the Wikidata community and citizen scientists alike. The first use case for this data is to populate Wikipedia Gene Wiki infoboxes directly from Wikidata with the data integrated above. This enables immediate updates of the Gene Wiki infoboxes as soon as the data in Wikidata is modified. Although Gene Wiki pages are currently only on the English language version of Wikipedia, the multilingual nature of Wikidata allows for a usage of the data we imported in all 280 different language Wikipedias. Apart from the Gene Wiki infobox use case, a powerful SPARQL endpoint and up to date exporting functionality (e.g. JSON, XML) enable very convenient further use of the data by scientists.\n\nIn summary, we created a fully open and extensible data resource for human and mouse molecular biology and biochemistry data. This resource enriches all the Wikipedias with structured information and serves as a new linking hub for the biological semantic web.

Bioinformatics

Transcriptomic analysis of diplomonad parasites reveals a trans-spliced intron in a helicase gene in Giardia

Gene expression is the central preoccupation of molecular biology, thus newly discovered facets of gene expression are of great interest. Recently, ourselves and others reported that in the diplomonad protist Giardia lamblia, the coding regions of several mRNAs are produced by ligation of independent RNA species expressed from distinct genomic loci. Such trans-splicing of introns was found to affect nearly as many genes in this organism as does classical cis-splicing of introns. These findings raised questions about the incidence of intron trans-splicing both across the G. lamblia transcriptome and across diplomonad diversity, however a dearth of transcriptomic data at the time prohibited systematic study of these questions. Here, I leverage newly available transcriptomic data from G. lamblia and the related diplomonad Spironucleus salmonicida to search for trans-spliced introns. My computational pipeline recovers all four previously reported trans-spliced introns in G. lamblia, suggesting good sensitivity. Scrutiny of thousands of potential cases revealed only a single additional trans-spliced intron in G. lamblia, in the p68 helicase gene, and no cases in S. salmonicida. The p68 intron differs from the previously reported trans-spliced introns in its high degree of streamlining: the core features of G. lamblia trans-spliced introns closely packed together, revealing striking efficiency in the implementation of a seemingly inherently inefficient molecular mechanism. These results serve to circumscribe the role of trans-splicing both in terms of genes effected and taxonomically. Future work should focus on the molecular mechanisms, evolutionary origins and phenotypic implications of this intriguing phenomenon.

Genomics

Thanatotranscriptome: genes actively expressed after organismal death

A continuing enigma in the study of biological systems is what happens to highly ordered structures, far from equilibrium, when their regulatory systems suddenly become disabled. In life, genetic and epigenetic networks precisely coordinate the expression of genes -- but in death, it is not known if gene expression diminishes gradually or abruptly stops or if specific genes are involved. We investigated the unwinding of the clock by identifying upregulated genes, assessing their functions, and comparing their transcriptional profiles through postmortem time in two species, mouse and zebrafish. We found transcriptional abundance profiles of 1,063 genes were significantly changed after death of healthy adult animals in a time series spanning from life to 48 or 96 h postmortem. Ordination plots revealed non-random patterns in profiles by time. While most thanatotranscriptome (thanatos-, Greek defn. death) transcript levels increased within 0.5 h postmortem, some increased only at 24 and 48 h. Functional characterization of the most abundant transcripts revealed the following categories: stress, immunity, inflammation, apoptosis, transport, development, epigenetic regulation, and cancer. The increase of transcript abundance was presumably due to thermodynamic and kinetic controls encountered such as the activation of epigenetic modification genes responsible for unraveling the nucleosomes, which enabled transcription of previously silenced genes (e.g., development genes). The fact that new molecules were synthesized at 48 to 96 h postmortem suggests sufficient energy and resources to maintain self-organizing processes. A step-wise shutdown occurs in organismal death that is manifested by the apparent upregulation of genes with various abundance maxima and durations. The results are of significance to transplantology and molecular biology.

Systems Biology

NCBI BLAST+ integrated into Galaxy

BackgroundThe NCBI BLAST suite has become ubiquitous in modern molecular biology, used for small tasks like checking capillary sequencing results of single PCR products through to genome annotation or even larger scale pan-genome analyses. For early adopters of the Galaxy web-based biomedical data analysis platform, integrating BLAST was a natural step for sequence comparison workflows.\n\nFindingsThe command line NCBI BLAST+ tool suite was wrapped for use within Galaxy, defining appropriate datatypes as needed, with the goal of making common BLAST tasks easy, and advanced tasks possible.\n\nConclusionsThis effort has been come an informal international collaborative effort, and is deployed and used on Galaxy servers worldwide. Several example use-cases are described herein.

Bioinformatics

Tying down loose ends in the Chlamydomonas genome

The Chlamydomonas genome has been sequenced, assembled and annotated to produce a rich resource for genetics and molecular biology in this well-studied model organism. The annotated genome is very rich in open reading frames upstream of the annotated coding sequence ( uORFs): almost three quarters of the assigned transcripts have at least one uORF, and frequently more than one. This is problematic with respect to the standard scanning model for eukaryotic translation initiation. These uORFs can be grouped into three classes: class 1, initiating in-frame with the coding sequence (cds) (thus providing a potential in-frame N-terminal extension); class 2, initiating in the 5UT and terminating out-of-frame in the cds; and class 3, initiating and terminating within the 5UT. Multiple bioinformatics criteria (including analysis of Kozak consensus sequence agreement and BLASTP comparisons to the closely related Volvox genome, and statistical comparison to cds and to random-sequence controls) indicate that of ~4000 class 1 uORFs, approximately half are likely in vivo translation initiation sites. The proposed resulting N-terminal extensions in many cases will sharply alter the predicted biochemical properties of the encoded proteins. These results suggest significant modifications in ~2000 of the ~20,000 transcript models with respect to translation initiation and encoded peptides. In contrast, class 2 uORFs may be subject to purifying selection, and the existent ones (surviving selection) are likely inefficiently translated. Class 3 uORFs are remarkably similar to random sequence expectations with respect to size, number and composition and therefore may be largely selectively neutral; their very high abundance (found in more than half of transcripts, frequently with multiple uORFs per transcript) nevertheless suggests the possibility of translational regulation on a wide scale.

Genomics

Patching holes in the Chlamydomonas genome

The Chlamydomonas genome has been sequenced, assembled and annotated to produce a rich resource for genetics and molecular biology in this well-studied model organism. However, the current reference genome contains ~1000 blocks of unknown sequence ( N-islands), which are frequently placed in introns of annotated gene models. We developed a strategy, using careful bioinformatics analysis of short-sequence cDNA and genomic DNA reads, to search for previously unknown exons hidden within such blocks, and determine the sequence and exon/intron boundaries of such exons. These methods are based on assembly and alignment completely independent of prior reference assembly or reference annotation. Our evidence indicates that ~one-quarter of the annotated intronic N-islands actually contain hidden exons. For most of these our algorithm recovers full exonic sequence with associated splice junctions and exon-adjacent intron sequence, that can be joined to the reference genome assembly and annotated transcript models. These new exons represent de novo sequence generally present nowhere in the assembled genome, and the added sequence can be shown in many cases to greatly improve evolutionary conservation of the predicted encoded peptides. At the same time, our results confirm the purely intronic status for a substantial majority of N-islands annotated as intronic in the reference annotated genome, increasing confidence in this valuable resource.

Genomics

Expanding perspectives on cognition in humans, animals, and machines

Over the past decade neuroscience has been attacking the problem of cognition with increasing vigor. Yet, what exactly is cognition, beyond a general signifier of anything seemingly complex the brain does? Here, we briefly review attempts to define, describe, explain, build, enhance and experience cognition. We highlight perspectives including psychology, molecular biology, computation, dynamical systems, machine learning, behavior and phenomenology. This survey of the landscape reveals not a clear target for explanation but a pluralistic and evolving scene with diverse opportunities for grounding future research. We argue that rather than getting to the bottom of it, over the next century, by deconstructing and redefining cognition, neuroscience will and should expand rather than merely reduce our concept of the mind.

Animal Behavior and Cognition

Statistical inference of protein structural alignments using information and compression

Structural molecular biology depends crucially on computational techniques that compare protein three-dimensional structures and generate structural alignments (the assignment of one-to-one correspondences between subsets of amino acids based on atomic coordinates.) Despite its importance, the structural alignment problem has not been formulated, much less solved, in a consistent and reliable way. To overcome these difficulties, we present here a framework for precise inference of structural alignments, built on the Bayesian and information-theoretic principle of Minimum Message Length (MML). The quality of any alignment is measured by its explanatory power - the amount of lossless compression achieved to explain the protein coordinates using that alignment. We have implemented this approach in the program MMLigner http://lcb.infotech.monash.edu.au/mmligner to distinguish statistically significant alignments, not available elsewhere. We also demonstrate the reliability of MMLigners alignment results compared with the state of the art. Importantly, MMLigner can also discover different structural alignments of comparable quality, a challenging problem for oligomers and protein complexes.

Bioinformatics

Genome-wide association study implicates immune activation of multiple integrin genes in inflammatory bowel disease

Genetic association studies have identified 210 risk loci for inflammatory bowel disease, which have revealed fundamental aspects of the molecular biology of the disease, including the roles of autophagy and Th17 cell signaling and development. We performed a genome-wide association study of 25,305 individuals, and meta-analyzed with published summary statistics, yielding a total sample size of 59,957 subjects. We identified 26 new genome-wide significant loci, three of which contain integrin genes that encode molecules in pathways identified as important therapeutic targets in inflammatory bowel disease. The associated variants are also correlated with expression changes in response to immune stimulus at two of these genes (ITGA4, ITGB8) and at two previously implicated integrin loci (ITGAL, ICAM1). In all four cases, the stimulus-dependent expression increasing allele also increases disease risk. We applied summary statistic fine-mapping and identified likely causal missense variants in the primary immune deficiency gene PLCG2 and the negative regulator of inflammation, SLAMF8. Our results demonstrate that new common variant associations continue to identify genes and pathways of relevance to therapeutic target identification and prioritization.

Genomics

Vibrio natriegens, a new genomic powerhouse

Recombinant DNA technology has revolutionized biomedical research with continual innovations advancing the speed and throughput of molecular biology. Nearly all these tools, however, are reliant on Escherichia coli as a host organism, and its lengthy growth rate increasingly dominates experimental time. Here we report the development of Vibrio natriegens, a free-living bacteria with the fastest generation time known, into a genetically tractable host organism. We systematically characterize its growth properties to establish basic laboratory culturing conditions. We provide the first complete Vibrio natriegens genome, consisting of two chromosomes of 3,248,023 bp and 1,927,310 bp that together encode 4,578 open reading frames. We reveal genetic tools and techniques for working with Vibrio natriegens. These foundational resources will usher in an era of advanced genomics to accelerate biological, biotechnological, and medical discoveries.

Genomics

Measurements of translation initiation from all 64 codons in E. coli

Our understanding of translation is one cornerstone of molecular biology that underpins our capacity to engineer living matter. The canonical start codon (AUG) and a few near-cognates (GUG, UUG) are typically considered as the \"start codons\" for translation initiation in Escherichia coli (E. coli). Translation is typically not thought to initiate from the 61 remaining codons. Here, we systematically quantified translation initiation in E. coli from all 64 triplet codons. We detected protein synthesis above background initiating from at least 46 codons. Translation initiated from these non-canonical start codons at levels ranging from 0.01% to 2% relative to AUG. Translation initiation from non-canonical start codons may contribute to the synthesis of peptides in both natural and synthetic biological systems

Synthetic Biology

Morphological plant modeling: Unleashing geometric and topologic potential within the plant sciences

Plant morphology is inherently mathematical in that morphology describes plant form and architecture with geometrical and topological descriptors. The geometries and topologies of leaves, flowers, roots, shoots and their spatial arrangements have fascinated plant biologists and mathematicians alike. Beyond providing aesthetic inspiration, quantifying plant morphology has become pressing in an era of climate change and a growing human population. Modifying plant morphology, through molecular biology and breeding, aided by a mathematical perspective, is critical to improving agriculture, and the monitoring of ecosystems with fewer natural resources. In this white paper, we begin with an overview of the mathematical models applied to quantify patterning in plants. We then explore fundamental challenges that remain unanswered concerning plant morphology, from the barriers preventing the prediction of phenotype from genotype to modeling the movement of leafs in air streams. We end with a discussion concerning the incorporation of plant morphology into educational programs. This strategy focuses on synthesizing biological and mathematical approaches and ways to facilitate research advances through outreach, cross-disciplinary training, and open science. This white paper arose from bringing mathematicians and biologists together at the National Institute for Mathematical and Biological Synthesis (NIMBioS) workshop titled \"Morphological Plant Modeling: Unleashing Geometric and Topological Potential within the Plant Sciences\" held at the University of Tennessee, Knoxville in September, 2015. Never has the need to quantify plant morphology been more imperative. Unleashing the potential of geometric and topological approaches in the plant sciences promises to transform our understanding of both plants and mathematics.

Plant Biology

CASTOR: A machine learning platform for reproducible viral genome classification

MotivationAdvances in cloning and sequencing technology yielded a massive number of genome of virus strains. The classification and annotation of these genomes constitute important assets in the discovery of genomic variability, taxonomic characteristics and disease mechanisms. Existing classification methods are often designed for a well-studied virus. Thus, the viral comparative genomic studies could benefit from more generic, fast and accurate tools for classifying and typing newly sequenced strains of diverse virus families.\n\nResultsHere, we introduce a fast, accurate and generic virus classification platform, CASTOR, based on a machine learning approach. CASTOR is inspired by a well-known technique in molecular biology: Restriction Fragment Length Polymorphism (RFLP). It simulates the restriction digestion of genomic material by different enzymes into fragments in-silico. It uses two metrics to construct feature vectors for machine learning algorithms in the classification step. We benchmark CASTOR for the classification of distinct datasets of Human Papillomaviruses (HPV), Hepatitis B Viruses (HBV) and Human Immunodeficiency viruses (HIV). Results reveal true positive rates of 99%, 99% and 98% for HPV Alpha species, HBV genotyping and HIV M group subtyping respectively. Furthermore, CASTOR shows a competitive performance compare to well-known HIV-specific classifier REGA and COMET on whole genome and pol fragments. With such prediction rates, genericity and robustness, as well as rapidity, such approach could constitute a reference in large-scale virus studies. Finally, we developed the CASTOR web platform for open access and reproducible viral machine learning classifiers.\n\nAvailabilityhttp://castor.bioinfo.uqam.ca\n\nContactdiallo.abdoulaye@uqam.ca

bioinformatics

Uniform Resolution of Compact Identifiers for Biomedical Data

Most biomedical data repositories issue locally-unique accessions numbers, but do not provide globally unique, machine-resolvable, persistent identifiers for their datasets, as required by publishers wishing to implement data citation in accordance with widely accepted principles. Local accessions may however be prefixed with a namespace identifier, providing global uniqueness. Such \"compact identifiers\" have been widely used in biomedical informatics to support global resource identification with local identifier assignment.\n\nWe report here on our project to provide robust support for machine-resolvable, persistent compact identifiers in biomedical data citation, by harmonizing the Identifiers.org and N2T.net (Name-To-Thing) meta-resolvers and extending their capabilities. Identifiers.org services hosted at the European Molecular Biology Laboratory - European Bioinformatics Institute (EMBL-EBI), and N2T.net services hosted at the California Digital Library (CDL), can now resolve any given identifier from over 600 source databases to its original source on the Web, using a common registry of prefix-based redirection rules.\n\nWe believe these services will be of significant help to publishers and others implementing persistent, machine-resolvable citation of research data.

bioinformatics

An extensible ontology for inference of emergent whole cell function from relationships between subcellular processes

Whole cell responses arise from coordinated interactions between diverse human gene products functioning within various pathways underlying sub-cellular processes (SCP). Lower level SCPs interact to form higher level SCPs, often in a context specific manner to give rise to whole cell function. We sought to determine if capturing such relationships enables us to describe the emergence of whole cell functions from interacting SCPs. We developed the \"Molecular Biology of the Cell\" ontology based on standard cell biology and biochemistry textbooks and review articles. Currently, our ontology contains 5,392 genes, 753 SCPs and 19,182 expertly curated gene-SCP associations. Our algorithm to populate the SCPs with genes enables extension of the ontology on demand and the adaption of the ontology to the continuously growing cell biological knowledge. Since whole cell responses most often arise from the coordinated activity of multiple SCPs, we developed a dynamic enrichment algorithm that flexibly predicts SCP-SCP relationships beyond the current taxonomy. This algorithm enables us to identify interactions between SCPs as a basis for higher order function in a context dependent manner, allowing us to provide a detailed description of how SCPs together can give rise to whole cell functions. We conclude that this ontology can, from omics data sets, enable the development of detailed multidimensional SCP networks for predictive modeling of emergent whole cell functions.

systems biology

Combining Bayesian Approaches and Evolutionary Techniques for the Inference of Breast Cancer Networks

Gene and protein networks are very important to model complex large-scale systems in molecular biology. Inferring or reverseengineering such networks can be defined as the process of identifying gene/protein interactions from experimental data through computational analysis. However, this task is typically complicated by the enormously large scale of the unknowns in a rather small sample size. Furthermore, when the goal is to study causal relationships within the network, tools capable of overcoming the limitations of correlation networks are required. In this work, we make use of Bayesian Graphical Models to attach this problem and, specifically, we perform a comparative study of different state-of-the-art heuristics, analyzing their performance in inferring the structure of the Bayesian Network from breast cancer data.

bioinformatics