bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Intra-protein binding peptide fragments have specific and intrinsic sequence patterns

The key finding in the DNA double helix model is the specific pairing or binding between nucleotides A-T and C-G, and the pairing rules are the molecule basis of genetic code. Unfortunately, no such rules have been discovered for proteins. Here we show that similar rules and intrinsic sequence patterns between intra-protein binding peptide fragments do exist, and they can be extracted using a deep learning algorithm. Multi-millions of binding and non-binding peptide fragments from currently available protein X-ray structures are classified with an accuracy of up to 93%. This discovery has the potential in helping solve protein folding and protein-protein interaction problems, two open and fundamental problems in molecular biology.\n\nOne Sentence SummaryClassification of binding and non-binding intra-protein peptide fragments using feed-forward neural network

bioinformatics

Interdependence, Reflexivity, Fidelity, Impedance Matching, And The Evolution Of Genetic Coding

Genetic coding is generally thought to have required ribozymes whose functions were taken over by polypeptide aminoacyl-tRNA synthetases (aaRS). Two discoveries about aaRS and their tRNA substrates now furnish a unifying rationale for the opposite conclusion: that the key processes of the Central Dogma of molecular biology emerged simultaneously and naturally from simple origins in a peptide*RNA partnership, eliminating the epistemological need for a prior RNA world. First, the two aaRS classes likely arose from opposite strands of the same ancestral gene, implying a simple genetic alphabet. Inversion symmetries in aaRS structural biology arising from genetic complementarity would have stabilized the initial and subsequent differentiation of coding specificities and hence rapidly promoted diversity in the proteome. Second, amino acid physical chemistry maps onto tRNA identity elements, establishing reflexivity in protein aaRS. Bootstrapping of increasingly detailed coding is thus intrinsic to polypeptide aaRS, but impossible in an RNA world. These notions underline the following concepts that contradict gradual replacement of ribozymal aaRS by polypeptide aaRS: (i) any set of aaRS must be interdependent; (ii) reflexivity intrinsic to polypeptide aaRS production dynamics promotes bootstrapping; (iii) takeover of RNA-catalyzed aminoacylation by enzymes will necessarily degrade specificity; (iv) the Central Dogmas emergence is most probable when replication and translation error rates remain comparable. These characteristics are necessary and sufficient for the essentially de novo emergence of a coupled gene-replicase-translatase system of genetic coding that would have continuously preserved the functional meaning of genetically encoded protein genes whose phylogenetic relationships match those observed today.

evolutionary biology

Insuperable Problems Of The Genetic Code Initially Emerging In An RNA World

Differential equations for error-prone information transfer (template replication, transcription or translation) are developed in order to consider, within the theory of autocatalysis, the advent of coded protein synthesis. Variations of these equations furnish a basis for comparing the plausibility of contrasting scenarios for the emergence of tRNA aminoacylation, ultimately by enzymes, and the relationship of this process with the origin of the universal system of molecular biological information processing embodied in the Central Dogma. The hypothetical RNA World does not furnish an adequate basis for explaining how this system came into being, but principles of self-organisation that transcend Darwinian natural selection furnish an unexpectedly robust basis for a rapid, concerted transition to genetic coding from a peptide*RNA world.

evolutionary biology

A Varroa Destructor Protein Atlas Reveals Molecular Underpinnings Of Developmental Transitions And Sexual Differentiation

Varroa destructor is the most economically damaging honey bee pest, weakening colonies by simultaneously parasitizing bees and transmitting harmful viruses. Despite these impacts on honey bee health, surprisingly little is known about its fundamental molecular biology. Here we present a Varroa protein atlas crossing all major developmental stages (egg, protonymph, deutonymph and adult) for both male and female mites as a web-based interactive tool (http://foster.nce.ubc.ca/varroa/index.html). By intensity-based label-free quantitation, 1,433 proteins were differentially expressed across developmental stages, including two distinct viral polyproteins. Enzymes for processing carbohydrates and amino acids were among many of these differences as well as proteins involved in cuticle formation. Lipid transport involving vitellogenin was the most significantly enriched biological process in the foundress (reproductive female) and young mites. In addition, we found that 101 proteins were sexually regulated and functional enrichment analysis suggests that chromatin remodeling may be a key feature of sex determination. In a proteogenomic effort, we identified 519 protein-coding regions (169 of which were differentially expressed) supported by 1,464 peptides which were previously unannotated. Since this is a recurring trend with annotating genomes of non-model species, we analyzed their amino acid and nucleotide composition as well as their orthology to other species to suggest reasons why they may have been missed initially. Overall, this work provides a first-of-its-kind interrogation of the patterns of gene expression that govern the Varroa life cycle and the tools we have developed will support further research on this threatening honey bee pest.

developmental biology

Simple, single-step, and scar-free mutagenesis of bacterial genes

The need for generating precisely designed mutations is common in genetics, biochemistry, and molecular biology. Here, I describe a new {lambda} Red recombineering method (Direct and Inverted Repeat stimulated excision; DIRex) for fast and easy generation of single point mutations, small insertions or replacements as well as deletions of any size, in bacterial genes. The method does not leave any resistance marker or scar sequence and requires only one transformation to generate a semi-stable intermediate insertion mutant. Spontaneous excision of the intermediate efficiently and accurately generates the final mutant. In addition, the intermediate is transferable between strains by generalized transductions, enabling transfer of the mutation into multiple strains without repeating the recombineering step. Existing methods that can be used to accomplish similar results are either (i) more complicated to design, (ii) more limited in what mutation types can be made, or (iii) require expression of extrinsic factors in addition to {lambda} Red. I demonstrate the utility of the method by generating several deletions, small insertions/replacements, and single nucleotide exchanges in Escherichia coli and Salmonella enterica. Furthermore, the design parameters that influence the excision frequency and the success rate of generating desired point mutations have been examined to determine design guidelines for optimal efficiency.

synthetic biology

Producing Hfq/Sm Proteins and sRNAs for Structural and Biophysical Studies of Ribonucleoprotein Assembly

Hfq is a bacterial RNA-binding protein that plays key roles in the post-transcriptional regulation of gene expression. Like other Sm proteins, Hfq assembles into toroidal discs that bind RNAs with varying affinities and degrees of sequence specificity. By simultaneously binding to a regulatory small RNA (sRNA) and an mRNA target, Hfq hexamers facilitate productive RNA[···]RNA interactions; the generic nature of this chaperone-like functionality makes Hfq a hub in many sRNA-based regulatory networks. That Hfq is crucial in diverse cellular pathways--including stress response, quorum sensing and biofilm formation-- has motivated genetic and RNAomic studies of its function and physiology (in vivo), as well as biochemical and structural analyses of Hfq[···]RNA interactions (in vitro). Indeed, crystallographic and bio-physical studies first established Hfq as a member of the phylogenetically-conserved Sm superfamily. Crystallography and other biophysical methodologies enable the RNA-binding properties of Hfq to be elucidated in atomic detail, but such approaches have stringent sample requirements, viz.: reconstituting and characterizing an Hfq*RNA complex requires ample quantities of well-behaved (sufficient purity, homogeneity) specimens of Hfq and RNA (sRNA, mRNA fragments, short oligoribonucleotides, or even single nucleotides). The production of such materials is covered in this Chapter, with a particular focus on recombinant Hfq proteins for crystallization experiments.\n\nAbbreviations\n\nJournal formatMethods in Molecular Biology (Springer Protocols series); this volume is entitled \"Bacterial Regulatory RNA: Methods and Protocols\"; an author guide is linked at http://www.springer.com/series/7651

biochemistry

Beyond Homology Transfer: Deep Learning for Automated Annotation of Proteins

Accurate annotation of protein functions is important for a profound understanding of molecular biology. A large number of proteins remain uncharacterized because of the sparsity of available supporting information. For a large set of uncharacterized proteins, the only type of information available is their amino acid sequence. In this paper, we propose DeepSeq - a deep learning architecture - that utilizes only the protein sequence information to predict its associated functions. The prediction process does not require handcrafted features; rather, the architecture automatically extracts representations from the input sequence data. Results of our experiments with DeepSeq indicate significant improvements in terms of prediction accuracy when compared with other sequence-based methods. Our deep learning model achieves an overall validation accuracy of 86.72%, with an F1 score of 71.13%. Moreover, using the automatically learned features and without any changes to DeepSeq, we successfully solved a different problem i.e. protein function localization, with no human intervention. Finally, we discuss how this same architecture can be used to solve even more complicated problems such as prediction of 2D and 3D structure as well as protein-protein interactions.

bioinformatics

Flexible, multi-use, PCR-based nucleic acid integrity assays based on the ubiquitin C gene

Nucleic acid integrity assessment is an important aspect of quality control for many applications in molecular biology. A number of methods exist (electrophoresis- or PCR-based), but they are not universally applicable. Some of them need huge amounts of input, process certain amount of samples, require expensive equipment, only work specifically on (c)DNA, (m)RNA, certain species or certain tissues, or produce fragments covering a small length range. We investigated if the ubiquitin C gene (UBC) could be used to develop flexible, multi-use PCR-based (deoxy)ribonucleic acid integrity assays. UBC gene analysis (in human, mouse, pig, cow, horse, sheep, dog and cat) shows that UBC is a highly conserved and ubiquitously expressed gene (reference gene in RT-qPCR), that encodes a polyubiquitin precursor (containing tandem repeats of at least 5 ubiquitin monomers of 228 bp) in a single exon. On average, ubiquitin monomers show a nucleic acid sequence identity of 96% at intraspecies level and 93% at interspecies level. Based on a multiple alignment of all monomer ubiquitin sequences of all investigated species, we could design a single degenerated primer pair generating PCR amplicons of 137, 365, 593 and 821 bp on low amounts of high quality DNA of all investigated species (down to 10 pg) and on cDNA reverse transcribed from high quality RNA from different tissues (e.g. heart, liver, brain, kidney). Increasing levels of nucleic acid degradation resulted in a decrease of amplification products starting with the longer amplicons. We conclude that UBC is suited to develop a single, cheap, universal assay to estimate the presence, integrity and amplificability of native mammalian nucleic acids. In addition, we used the same strategy to design a similar assay to check the quality of bisulfite treated mammalian DNA.

genetics

Biosensor libraries harness large classes of binding domains for allosteric transcription regulators

Bacterias ability to specifically sense small molecules in their environment and trigger metabolic responses in accordance is an invaluable biotechnological resource. While many transcription factors (TFs) mediating these processes have been studied, only a handful has been leveraged for molecular biology applications. To expand this panel of biotechnologically important sensors here we present a strategy for the construction and testing of chimeric TF libraries, based on the fusion of highly soluble periplasmic binding proteins (PBPs) with DNA-binding domains (DBDs). We validated this strategy by constructing and functionally testing two unique sense-and-response regulators for benzoate, an environmentally and industrially relevant metabolite. This work will enable the development of tailored biosensors for synthetic regulatory circuits.

synthetic biology

A toolbox of anti-mouse and rabbit IgG secondary nanobodies

Polyclonal anti-IgG secondary antibodies are essential tools for many molecular biology techniques and diagnostic tests. Their animal-based production is, however, a major ethical problem. Here, we introduce a sustainable alternative, namely nanobodies against all mouse IgG subclasses and rabbit IgG. They can be produced at large scale in E. coli and could thus make secondary antibody-production in animals obsolete. Their recombinant nature allows fusion with affinity tags or reporter enzymes as well as efficient maleimide chemistry for fluorophore-coupling. We demonstrate their superior performance in Western Blotting, both in peroxidase- and fluorophore-linked form. Their site-specific labeling with multiple fluorophores creates bright imaging reagents for confocal and super-resolution microscopy with much smaller label displacement than traditional secondary antibodies. They also enable simpler and faster immunostaining protocols and even allow multi-target localization with primary IgGs from the same species and of the same class.

biochemistry

MAPLE: a Modular Automated Platform for Large-scale Experiments, a low-cost robot for integrated animal-handling and phenotyping

Genetic model system animals have significant scientific value in part because of large-scale experiments like screens, but performing such experiments over long time periods by hand is arduous and risks errors. Thus the field is poised to benefit from automation, just as molecular biology did from liquid-handling robots. We developed a Modular Automated Platform for Large-scale Experiments (MAPLE), a Drosophila-handling robot capable of conducting lab tasks and experiments. We demonstrate MAPLEs ability to accelerate the collection of virgin female flies (a pervasive experimental chore in fly genetics) and assist high-throughput phenotyping assays. Using MAPLE to autonomously run a novel social interaction experiment, we found that 1) pairs of flies exhibit persistent idiosyncrasies in affiliative behavior, 2) these dyad-specific interactions require olfactory and visual cues, and 3) social interaction network structure is topologically stable over time. These diverse examples demonstrate MAPLEs versatility as a general platform for conducting fly science automatically.

bioengineering

Time- and cost-efficient high-throughput transcriptomics enabled by Bulk RNA Barcoding and sequencing

Genome-wide gene expression analyses by RNA sequencing (RNA-seq) have quickly become a standard in molecular biology because of the widespread availability of high throughput sequencing technologies. While powerful, RNA-seq still has several limitations, including the time and cost of library preparation, which makes it difficult to profile many samples simultaneously. To deal with these constraints, the single-cell transcriptomics field has implemented the early multiplexing principle, making the library preparation of hundreds of samples (cells) markedly more affordable. However, the current standard methods for bulk transcriptomics (such as TruSeq Stranded mRNA) remain expensive, and relatively little effort has been invested to develop cheaper, but equally robust methods. Here, we present a novel approach, Bulk RNA Barcoding and sequencing (BRB-seq), that combines the multiplexing-driven cost-effectiveness of a single-cell RNA-seq workflow with the performance of a bulk RNA-seq procedure. BRB-seq produces 3 enriched cDNA libraries that exhibit similar gene expression quantification to TruSeq and that maintain this quality, also in terms of number of detected differentially expressed genes, even with low quality RNA samples. We show that BRB-seq is about 25 times less expensive than TruSeq, enabling the generation of ready to sequence libraries for up to 192 samples in a day with only 2 hours of hands-on time. We conclude that BRB-seq constitutes a powerful alternative to TruSeq as a standard bulk RNA-seq approach. Moreover, we anticipate that this novel method will eventually replace RT-qPCR-based gene expression screens given its capacity to generate genome-wide transcriptomic data at a cost that is comparable to profiling 4 genes using RT-qPCR.\n\n SoftwareWe developed a suite of open source tools (BRB-seqTools) to aid with processing BRB-seq data and generating count matrices that are used for further analyses. This suite can perform demultiplexing, generate count/UMI matrices and trim BRB-seq constructs and is freely available at http://github.com/DeplanckeLab/BRB-seqTools\n\nHighlightsO_LIRapid (~2h hands on time) and low-cost approach to perform transcriptomics on hundreds of RNA samples\nC_LIO_LIStrand specificity preserved\nC_LIO_LIPerformance: number of detected genes is equal to Illumina TruSeq Stranded mRNA at same sequencing depth\nC_LIO_LIHigh capacity: low cost allows increasing the number of biological replicates\nC_LIO_LIProduces reliable data even with low quality RNA samples (down to RIN value = 2)\nC_LIO_LIComplete user-friendly sequencing data pre-processing and analysis pipeline allowing result acquisition in a day\nC_LI

genomics

Engineering of new-to-nature halogenated indigo precursors in plants

Plants are versatile chemists producing a tremendous variety of specialized compounds. Here, we describe the engineering of entirely novel metabolic pathways in planta enabling generation of halogenated indigo precursors as non-natural plant products.\n\nIndican (indolyl-{beta}-D-glucopyranoside) is a secondary metabolite characteristic of a number of dyers plants. Its deglucosylation and subsequent oxidative dimerization leads to the blue dye, indigo. Halogenated indican derivatives are commonly used as detection reagents in histochemical and molecular biology applications; their production, however, relies largely on chemical synthesis. To attain the de novo biosynthesis in a plant-based system devoid of indican, we employed a sequence of enzymes from diverse sources, including three microbial tryptophan halogenases substituting the amino acid at either C5, C6, or C7 of the indole moiety. Subsequent processing of the halotryptophan by bacterial tryptophanase TnaA in concert with a mutant of the human cytochrome P450 monooxygenase 2A6 and glycosylation of the resulting indoxyl derivatives by an endogenous tobacco glucosyltransferase yielded corresponding haloindican variants in transiently transformed Nicotiana benthamiana plants. Accumulation levels were highest when the 5-halogenase PyrH was utilized, reaching 0.93 {+/-}0.089 mg/g dry weight of 5-chloroindican. The identity of the latter was unambiguously confirmed by NMR analysis. Moreover, our combinatorial approach, facilitated by the modular assembly capabilities of the GoldenBraid cloning system and inspired by the unique compartmentation of plant cells, afforded testing a number of alternative subcellular localizations for pathway design. In consequence, chloroplasts were validated as functional biosynthetic venues for haloindican, with the requisite reducing augmentation of the halogenases as well as the cytochrome P450 monooxygenase fulfilled by catalytic systems native to the organelle.\n\nThus, our study puts forward a viable alternative production platform for halogenated fine chemicals, eschewing reliance on fossil fuel resources and toxic chemicals. We further contend that in planta generation of halogenated indigoid precursors previously unknown to nature offers an extended view on and, indeed, pushes forward the established frontiers of biosynthetic capacity of plants.\n\nGraphical abstract\n\nHighlightsO_LIEntirely novel combinatorial pathway yielded new-to-nature halogenated indigoids.\nC_LIO_LIAn array of specific chloro- and bromoindican regioisomers was generated.\nC_LIO_LISignificant retrieval rates afforded unambiguous metabolite identification.\nC_LIO_LIEngineered plants emerge as alternative manufacture chassis for fine chemicals.\nC_LI

synthetic biology

CONSTRUCTION OF WHOLE GENOMES FROM SCAFFOLDS USING SINGLE CELL STRAND-SEQ DATA

Accurate reference genome sequences provide the foundation for modern molecular biology and genomics as the interpretation of sequence data to study evolution, gene expression and epigenetics depends heavily on the quality of the genome assembly used for its alignment. Correctly organising sequenced fragments such as contigs and scaffolds in relation to each other is a critical and often challenging step in the construction of robust genome references. We previously identified misoriented regions in the mouse and human reference assemblies using Strand-seq, a single cell sequencing technique that preserves DNA directionality1, 2. Here we demonstrate the ability of Strand-seq to build and correct full-length chromosomes, by identifying which scaffolds belong to the same chromosome and determining their correct order and orientation, without the need for overlapping sequences. We demonstrate that Strand-seq exquisitely maps assembly fragments into large related groups and chromosome-sized clusters without using new assembly data. Using template strand inheritance as a bi-allelic marker, we employ genetic mapping principles to cluster scaffolds that are derived from the same chromosome and order them within the chromosome based solely on directionality of DNA strand inheritance. We prove the utility of our approach by generating improved genome assemblies for several model organisms including the ferret, pig, Xenopus, zebrafish, Tasmanian devil and the Guinea pig.

genomics

AIControl: Replacing matched control experiments with machine learning improves ChIP-seq peak identification

ChIP-seq is a technique to determine binding locations of transcription factors, which remains a central challenge in molecular biology. Current practice is to use a \"control\" dataset to remove background signals from a immunoprecipitation (IP) target dataset. We introduce the AlControl framework, which eliminates the need to obtain a control dataset and instead identifies binding peaks by estimating the distributions of background signals from many publicly available control ChIP-seq datasets. We thereby avoid the cost of running control experiments while simultaneously increasing the accuracy of binding location identification. Specifically, AIControl can (1) estimate background signals at fine resolution, (2) systematically weigh the most appropriate control datasets in a data-driven way, (3) capture sources of potential biases that may be missed by one control dataset, and (4) remove the need for costly and time-consuming control experiments. We applied AIControl to 410 IP datasets in the ENCODE ChIP-seq database, using 440 control datasets from 107 cell types to impute background signal. Without using matched control datasets, AIControl identified peaks that were more enriched for putative binding sites than those identified by other popular peak callers that used a matched control dataset. We also demonstrated that our framework identifies binding sites that recover documented protein interactions more accurately.

bioinformatics

Widespread transcriptional scanning in testes modulates gene evolution rates

A long-standing question in molecular biology relates to why the testes express the largest number of genes relative to all other organs. Here, we report a detailed gene expression map of human spermatogenesis using single-cell RNA-Seq. Surprisingly, we found that spermatogenesis-expressed genes contain significantly fewer germline mutations than unexpressed genes, with the lowest mutation rates on the transcribed DNA strands. These results suggest a model of transcriptional scanning to reduce germline mutations by correcting DNA damage. This model also explains the rapid evolution in sensory- and immune-defense related genes, as well as in male reproduction genes. Collectively, our results indicate that widespread expression in the testes achieves a dual mechanism for maintaining the DNA integrity of most genes, while selectively promoting variation of other genes.

evolutionary biology

An analysis of current state of the art software on nanopore metagenomic data

ContextOur insight into DNA is controlled through a process called sequencing. Until recently, it was only possible to sequence DNA into short strings called \"reads\". Nanopore is a new sequencing technology to produce significantly longer reads. Using nanopore sequencing, a single molecule of DNA can be sequenced without the need for time consuming PCR amplification (polymerase chain reaction is a technique used in molecular biology to amplify a single copy or a few copies of a segment of DNA across several orders of magnitude).\n\nAimsMetagenomics is the study of genetic material recovered from environmental samples. A research team from IBERS (Institute of Biological, Environmental & Rural Sciences) at Aberystwyth University have sampled metagenomes from a coal mine in South Wales using the Nanopore MinION and given initial taxonomic (classification of organisms) summaries of the contents of the microbial community.\n\nMethodsUsing various new software aimed for metagenomic data, we are interested to discover how well current bioinformatics software works with the data-set. We will conduct analysis and research into how well these new state of the art software works with this new long read data and try out some recent new developments for such analysis.\n\nResultsMost of the software we used worked very well: we gained understanding of the ACGT count and quality of the data. However some software for bioinformatics dont seem to work with nanopore data. Furthermore, we can conclude that low quality nanopore data may actually be quite average.

bioinformatics

Proteomic and Evolutionary Analyses of Sperm Activation Identify Uncharacterized Genes in Caenorhabditis Nematodes

BackgroundNematode sperm have unique and highly diverged morphology and molecular biology. In particular, nematode sperm contain subcellular vesicles known as membranous organelles that are necessary for male fertility, yet play a still unknown role in overall sperm function. Here we take a novel proteomic approach to characterize the functional protein complement of membranous organelles in two Caenorhabditis species: C. elegans and C. remanei.\n\nResultsWe identify distinct protein compositions between membranous organelles and the activated sperm body. Two particularly interesting and undescribed gene families--the Nematode-Specific Peptide family, group D and the here designated Nematode-Specific Peptide family, group F--localize to the membranous organelle. Both multigene families are nematode-specific and exhibit patterns of conserved evolution specific to the Caenorhabditis clade. These data suggest gene family dynamics may be a more prevalent mode of evolution than sequence divergence within sperm. Using a CRISPR-based knock-out of the NSPF gene family, we find no evidence of a male fertility effect of these genes, despite their high protein abundance within the membranous organelles.\n\nConclusionsOur study identifies key components of this unique subcellular sperm component and establishes a path toward revealing their underlying role in reproduction.

evolutionary biology