bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Wx: a neural network-based feature selection algorithm for next-generation sequencing data

MotivationNext-generation sequencing (NGS), which allows the simultaneous sequencing of billions of DNA fragments simultaneously, has revolutionized how we study genomics and molecular biology by generating genome-wide molecular maps of molecules of interest. For example, an NGS-based transcriptomic assay called RNA-seq can be used to estimate the abundance of approximately 190,000 transcripts together. As the cost of next-generation sequencing sharply declines, researchers in many fields have been conducting research using NGS. The amount of information produced by NGS has made it difficult for researchers to choose the optimal set of target genes (or genomic loci).\n\nResultsWe have sought to resolve this issue by developing a neural network-based feature (gene) selection algorithm called Wx. The Wx algorithm ranks genes based on the discriminative index (DI) score that represents the classification power for distinguishing given groups. With a gene list ranked by DI score, researchers can institutively select the optimal set of genes from the highest-ranking ones. We applied the Wx algorithm to a TCGA pan-cancer gene-expression cohort to identify an optimal set of gene-expression biomarker (universal gene-expression biomarkers) candidates that can distinguish cancer samples from normal samples for 12 different types of cancer. The 14 gene-expression biomarker candidates identified by Wx were comparable to or outperformed previously reported universal gene expression biomarkers, highlighting the usefulness of the Wx algorithm for next-generation sequencing data. Thus, we anticipate that the Wx algorithm can complement current state-of-the-art analytical applications for the identification of biomarker candidates as an alternative method.\n\nAvailabilityhttps://github.com/deargen/DearWX\n\nContactkangk1204@dankook.ac.kr\n\nSupplementary informationSupplementary data are available at online.

bioinformatics

Sequence and primer independent stochastic heterogeneity in PCR amplification efficiency revealed by single molecule barcoding

The polymerase chain reaction (PCR) is one of the most widely used techniques in molecular biology. In combination with High Throughput Sequencing (HTS), PCR is widely used to quantify transcript abundance for RNA-seq and especially in the context of analysis of T cell and B cell receptor repertoires. In this study, we combine molecular DNA barcoding with HTS to quantify PCR output from individual target molecules. Our results demonstrate that the PCR process exhibits very significant unexpected heterogeneity, which is independent of the sequence of the primers or target, and independent of bulk experimental conditions. The mechanistic origin of this heterogeneity is not clear, but simulations suggest that it must derive from inherited differences between different DNA molecules within the reaction. The results illustrate that single molecule barcoding is important in order to derive reproducible quantitative results from any protocol which combines PCR with HTS.

Molecular Biology

Detection of functional protein domains by unbiased genome-wide forward genetic screening

Genetic and chemo-genetic interactions have played key roles in elucidating the molecular mechanisms by which certain chemicals perturb cellular functions. Many studies have employed gene knockout collections or gene disruption/depletion strategies to identify routes for evolving resistance to chemical agents. By contrast, searching for point-mutational genetic suppressors that can identify separation- or gain-of-function mutations, has been limited even in simpler, genetically amenable organisms such as yeast, and has not until recently been possible in mammalian cell culture systems. Here, by demonstrating its utility in identifying suppressors of cellular sensitivity to the drugs camptothecin or olaparib, we describe an approach allowing systematic, large-scale detection of spontaneous or chemically-induced suppressor mutations in yeast and in haploid mouse embryonic stem cells in a short timeframe, and with potential applications in essentially any other haploid system. In addition to its utility for molecular biology research, this protocol can be used to identify drug targets and to predict mechanisms leading to drug resistance. Mapping suppressor mutations on the primary sequence or three-dimensional structures of protein suppressor hits provides insights into functionally relevant protein domains, advancing our molecular understanding of protein functions, and potentially helping to improve drug design and applicability.

molecular biology

Molecular tools for phytoplankton monitoring samples

HABs can have severe impacts in fisheries or human health by the consumption of contaminated bivalves. Monitoring assessment (quantitative and qualitative identification) of these organisms, is routinely accomplished by microscopic identification and counting of these organisms. Nonetheless, molecular biology techniques are gaining relevance, once these approaches can easily identify phytoplankton organisms at species level and even cell number quantifications. This work tests 12 methods/kits for genomic DNA extraction and seven DNA polymerases to determine which is the best method for routinely use in a common molecular laboratory, for phytoplankton monitoring samples analyses. From our work, Direct PCR master mix for tissue samples, proved to be the most adequate by its velocity of processivity, practicability, reproducibility, sensitiveness and robustness. However, brands such as Omega Biotek, GRISP, Qiagen and MP Biomedicals also showed good results for conventional DNA extraction as well as all the Taq brands tested (GRISP, GE Healthcare Life Sciences, ThermoFisher Scientific and Promega). Lugols solution, with our tested kits did not show negative interference in DNA amplification. The same can be said about mechanical digestion, with no significant differences among kits with or without this homogenization step.

molecular biology

Microbial contamination screening and interpretation for biological laboratory environments

Advances in microbiome researches have led us to the realization that the composition of microbial communities of indoor environment is profoundly affected by the function of buildings, and in turn may bring detrimental effects to the indoor environment and the occupants. Thus investigation is warranted for a deeper understanding of the potential impact of the indoor microbial communities. Among these environments, the biological laboratories stand out because they are relatively clean and yet are highly susceptible to microbial contaminants. In this study, we assessed the microbial compositions of samples from the surfaces of various sites across different types of biological laboratories. We have qualitatively and quantitatively assessed these possible microbial contaminants, and found distinct differences in their microbial community composition. We also found that the type of laboratories has a larger influence than the sampling site in shaping the microbial community, in terms of both structure and richness. On the other hand, the public areas of the different types of laboratories share very similar sets of microbes. Tracing the main sources of these microbes, we identified both environmental and human factors that are important factors in shaping the diversity and dynamics of these possible microbial contaminations in biological laboratories. These possible microbial contaminants that we have identified will be helpful for people who aim to eliminate them from samples.\n\nImportanceMicrobial communities from biological laboratories might hamper the conduction of molecular biology experiments, yet these possible contaminations are not yet carefully investigated. In this work, a metagenomic approach has been applied to identify the possible microbial contaminants and their sources, from the surfaces of various sites across different types of biological laboratories. We have found distinct differences in their microbial community compositions. We have also identified the main sources of these microbes, as well as important factors in shaping the diversity and dynamics of these possible microbial contaminations. The identification and interpretation of these possible microbial contaminants in biological laboratories would be helpful for alleviate their potential detrimental effects.

microbiology

Evolutionarily informed deep learning methods: Predicting transcript abundance from DNA sequence

Deep learning methodologies have revolutionized prediction in many fields, and show potential to do the same in molecular biology and genetics. However, applying these methods in their current forms ignores evolutionary dependencies within biological systems and can result in false positives and spurious conclusions. We developed two novel approaches that account for evolutionary relatedness in machine learning models: 1) gene-family guided splitting, and 2) ortholog contrasts. The first approach accounts for evolution by constraining the models training and testing sets to include different gene families. The second, uses evolutionarily informed comparisons between orthologous genes to both control for and leverage evolutionary divergence during the training process. The two approaches were explored and validated within the context of mRNA expression level prediction, and have prediction auROC values ranging from 0.72 to 0.94. Model weight inspections showed biologically interpretable patterns, resulting in the novel hypothesis that the 3 UTR is more important for fine tuning mRNA abundance levels while the 5 UTR is more important for large scale changes.

molecular biology

Minimizing carry-over PCR contamination in expanded CAG/CTG repeat instability applications

Expanded CAG/CTG repeats underlie the aetiology of 14 neurological and neuromuscular disorders. The size of the repeat tract determines in large part the severity of these disorders with longer tracts causing more severe phenotypes. Expanded CAG/CTG repeats are also unstable in somatic tissues, which is thought to modify disease progression. Routine molecular biology applications involving these repeats, including quantifying their instability, are plagued by low PCR yields. This leads to the need for setting up more PCRs of the same locus, thereby increasing the risk of carry-over contamination. Here we aimed to reduce this risk by pre-treating the samples with a Uracil N-Glycosylase (Ung) and using dUTP instead of dTTP in PCRs. We successfully applied this method to the PCR amplification of expanded CAG/CTG repeats, their sequencing, and their molecular cloning. In addition, we optimized the gold-standard method for measuring repeat instability, small-pool PCR, such that it can be used together with Ung and dUTP-containing PCRs, without compromising data quality. We expect that the protocols herein to be applicable for molecular diagnostics of expanded repeat disorders and to manipulate other tandem repeats.

molecular biology

Monitoring the 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U) modification in eukaryotic tRNAs via the gamma-toxin endonuclease

The post-transcriptional modification of tRNA at the wobble position is a universal process occurring in all domains of life. In eukaryotes, the wobble uridine of particular tRNAs is transformed to the 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U) modification which is critical for proper mRNA decoding and protein translation. However, current methods to detect mcm5s2U are technically challenging and/or require specialized instrumental expertise. Here, we show that gamma-toxin endonuclease from the yeast Kluyveromyces lactis can be used as a probe for assaying mcm5s2U status in the tRNA of diverse eukaryotic organisms ranging from protozoans to mammalian cells. The assay couples the mcm5s2U-dependent cleavage of tRNA by gamma-toxin with standard molecular biology techniques such as Northern blot analysis or quantitative PCR to monitor mcm5s2U levels in multiple tRNA isoacceptors. The results gained from the gamma-toxin assay reveals the evolutionary conservation of the mcm5s2U modification across eukaryotic species. Moreover, we have employed the gamma-toxin assay to verify uncharacterized eukaryotic Trm9 and Trm112 homologs that catalyze the formation of mcm5s2U. These findings demonstrate the use of gamma-toxin as a detection method to monitor mcm5s2U status in diverse eukaryotic cell types for cellular, genetic and biochemical studies.

molecular biology

Control mechanisms for stochastic biochemical systems via computation of reachable sets

Controlling the behaviour of cells by rationally guiding molecular processes is an overarching aim of much of synthetic biology. Molecular processes, however, are notoriously noisy and frequently non-linear. We present an approach to studying the impact of control measures on motifs of molecular interactions, that addresses the problems faced in biological systems: stochasticity, parameter uncertainty, and non-linearity. We show that our reachability analysis formalism can describe the potential behaviour of biological (naturally evolved as well as engineered) systems, and provides a set of bounds on their dynamics at the level of population statistics: for example, we can obtain the possible ranges of means and variances of mRNA and protein expression levels, even in the presence of uncertainty about model parameters.

Systems Biology

Attempts to implement CRISPR/Cas9 for genome editing in the oomycete Phytophthora infestans

Few techniques have revolutionized the molecular biology field as much as genome editing using CRISPR/Cas9. Recently, a CRISPR/Cas9 system has been developed for the oomycete Phytophthora sojae, and since then it has been employed in two other Phytophthora spp. Here, we report our progress on efforts to establish the system in the potato late blight pathogen Phytophthora infestans. Using the original constructs as developed for P. sojae, we did not obtain any transformants displaying a mutagenized target gene. We made several modifications to the CRISPR/Cas9 system to pinpoint the reason for failure and also explored the delivery of pre-assembled ribonucleoprotein complexes. With this report we summarize an extensive experimental effort pursuing the application of a CRISPR/Cas9 system for targeted mutagenesis in P. infestans and we conclude with suggestions for future directions.

molecular biology

Efficient and flexible strategies for gene cloning and vector construction using overlap PCR

Gene cloning and vector construction are basic technologies in modern molecular biology for gene functional study. Here, we present flexible and efficient strategies for gene cloning and vector construction using overlap PCR. We firstly cloned the open reading frames (ORFs) of the porcine MSTN, chicken OVA and human -glucosidase genes by overlap PCR-based assembling of their exons, which could be amplified with genomic DNAs as the templates without RNA extraction and RT-PCR reaction. Secondly, we generated additionally three designed functional cassettes by overlap PCR-based assembling of different DNA elements, which facilitated the construction their expression vectors greatly. Moreover, we further developed an interesting overlap-circled PCR method for fast plasmid vector construction without any cutting and ligating procedure. These advanced applications of overlap PCR provide useful alternative tools for gene cloning and vector construction.

molecular biology

Enzyme-based synthesis of single molecule RNA FISH probes

Single molecule RNA fluorescence in situ hybridization (smRNA FISH) allows the quantitative analysis of gene expression in single cells. The technique relies on the use of pools of end-labeled fluorescent oligonucleotides to detect specific cellular RNA sequences. These fluorescent probes are currently chemically synthesized. Here I describe a novel technique based on the use of routine molecular biology enzymes to generate smRNA FISH probes without the need for chemical synthesis of pools of oligonucleotides. The protocol comprises 3 main steps: purification of phagemid-derived single stranded DNA molecules comprising a segment complementary to a target RNA sequence; fragmentation of these molecules by limited DNase I digestion; and end-labeling of the resulting oligonucleotides with terminal deoxynucleotide transferase and fluorescent dideoxynucleotides. smRNA FISH probes that are obtained using the technique presented here are shown to perform as well as conventional probes. The main advantages of the method are the low cost of probes and the flexibility it affords in the choice of labels. Enzyme-based synthesis of probes should further increase the popularity of smRNA FISH as a tool to investigate gene expression at the cellular or subcellular level.

molecular biology

The administration of high-mobility group box 1 fragment prevents deterioration of cardiac performance by enhancement of bone marrow mesenchymal stem cell homing in the delta-sarcoglycan-deficient hamster

ObjectivesWe hypothesized that systemic administration of high-mobility group box 1 fragment attenuates the progression of myocardial fibrosis and cardiac dysfunction in a hamster model of dilated cardiomyopathy by recruiting bone marrow mesenchymal stem cells thus causing enhancement of a self-regeneration system.\n\nMethodsTwenty-week-old J2N-k hamsters, which are {delta}-sarcoglycan-deficient, were treated with systemic injection of high-mobility group box 1 fragment (HMGB1, n=15) or phosphate buffered saline (control, n=11). Echocardiography for left ventricular function, cardiac histology, and molecular biology were analyzed. The life-prolonging effect was assessed separately using the HMGB1 and control groups, in addition to a monthly HMGB1 group which received monthly systemic injections of high-mobility group box 1 fragment, 3 times (HMGB1, n=11, control, n=9, monthly HMGB1, n=9).\n\nResultsThe HMGB1 group showed improved left ventricular ejection fraction, reduced myocardial fibrosis, and increased capillary density. The number of platelet-derived growth factor receptor-alpha and CD106 positive mesenchymal stem cells detected in the myocardium was significantly increased, and intra-myocardial expression of tumor necrosis factor stimulating gene 6, hepatic growth factor, and vascular endothelial growth factor were significantly upregulated after high-mobility group box 1 fragment administration. Improved survival was observed in the monthly HMGB1 group compared with the control group.\n\nConclusionsSystemic high-mobility group box 1 fragment administration attenuates the progression of left ventricular remodeling in a hamster model of dilated cardiomyopathy by enhanced homing of bone marrow mesenchymal stem cells into damaged myocardium, suggesting that high-mobility group box 1 fragment could be a new treatment for dilated cardiomyopathy.

molecular biology

Development of a papillation assay using constitutive promoters to find hyperactive transposases

BackgroundTransposable elements (TEs) form a diverse group of DNA sequences encoding functions for their own mobility. This ability has been exploited as a powerful tool for molecular biology and genomics techniques. However, their use is sometimes limited because their activity is auto-regulated to allow them to cohabit within their hosts without causing excessive genomic damage. To overcome these limitations, it is important to develop efficient and simple screening assays for hyperactive transposases.\n\nResultsTo widen the range of transposase expression normally accessible with inducible promoters, we have constructed a set of vectors based on constitutive promoters of different strengths. We characterized and validated our expression vectors with Hsmar1, a member of the mariner transposon family. We observed the highest rate of transposition with the weakest promoters. We went on to investigate the effects of mutations in the Hsmar1 transposase dimer interface and of covalently linking two transposase monomers in a single-chain dimer. We also tested the severity of mutations in the lineage leading to the human SETMAR gene, in which one copy of the Hsmar1 transposase has contributed a domain.\n\nConclusionsWe generated a set of vectors to provide a wide range of transposase expression which will be useful for screening libraries of transposase mutants. We also found that mutations in the Hsmar1 dimer interface provides resistance to overproduction inhibition in bacteria, which could be valuable for improving bacterial transposon mutagenesis techniques.

molecular biology

Bio-On-Magnetic-Beads (BOMB): Open platform for high-throughput nucleic acid extraction and manipulation

Current molecular biology laboratories rely heavily on the purification and manipulation of nucleic acids. Yet, commonly used centrifuge-and column-based protocols require specialised equipment, often use toxic reagents and are not economically scalable or practical to use in a high-throughput manner. Although it has been known for some time that magnetic beads can provide an elegant answer to these issues, the development of open-source protocols based on beads has been limited. In this article, we provide step-by-step instructions for an easy synthesis of functionalised magnetic beads, and detailed protocols for their use in the high-throughput purification of plasmids, genomic DNA and total RNA from different sources, as well as environmental TNA and PCR amplicons. We also provide a bead-based protocol for bisulfite conversion, and size selection of DNA and RNA fragments. Comparison to other methods highlights the capability, versatility and extreme cost-effectiveness of using magnetic beads. These open source protocols and the associated webpage (https://bomb.bio) can serve as a platform for further protocol customisation and community engagement.

molecular biology

RNA-DEPENDENT SYNTHESIS OF MAMMALIAN mRNA: IDENTIFICATION OF CHIMERIC INTERMEDIATE AND PUTATIVE END PRODUCT

Our initial understanding of the flow of protein-encoding genetic information, DNA to RNA to protein, a process defined as the \"central dogma of molecular biology\", was eventually amended to account for the information back-flow from RNA to DNA (reverse transcription), and for its \"side-flow\" from RNA to RNA (RNA-dependent RNA synthesis, RdRs). These processes, both potentially leading to protein production, were described only in viral systems, and although putative RNA-dependent RNA polymerase (RdRp) was shown to be present, and RdRs to occur, in most, if not all, mammalian cells, its function was presumed to be restricted to regulatory. Here we report the occurrence of protein-encoding RNA to RNA information transfer in mammalian cells. We describe below the detection, by next generation sequencing (NGS), of a chimeric doublestranded/pinhead intermediate containing both sense and antisense globin RNA strands covalently joined in a predicted and uniquely defined manner, whose cleavage at the pinhead would result in the generation of an endproduct containing the intact coding region of the original mRNA. We also describe the identification of the putative end product of RNA-dependent globin mRNA amplification. It is heavily modified, uniformly truncated at both untranslated regions (UTRs), terminates with the OH group at the 5 end, consistent with a cleavagegenerated 5 terminus, and its massive cellular amount is unprecedented for a conventional mRNA transcription product. It also translates in a cell-free system into polypeptides indistinguishable from the translation product of conventional globin mRNA. The physiological significance of the mammalian mRNA amplification, which might operate during terminal differentiation and in the production of highly abundant rapidly generated proteins such as some collagens or other components of extracellular matrix, with every genome-originated mRNA molecule acting as a potential template, as well as possible implications, including physiologically occurring intracellular PCR process, iPCR, are discussed in the paper.

Molecular Biology

The expression tractability of a biological trait

Understanding how gene expression is translated to phenotype is central to modern molecular biology, but the success is contingent on the intrinsic tractability of the specific traits under examination. However, an a priori estimate of trait tractability from the perspective of gene expression is unavailable. Motivated by the concept of entropy in a thermodynamic system, we here propose such an estimate (ST) by gauging the number (N) of different expression states that underlie the same trait abnormality, with large ST corresponding to large N. By analyzing over 200 yeast morphological traits we show that ST is constrained by natural selection, which builds co-regulated gene modules to minimize the total number of possible expression states. We further show that ST is a good measure of the titer of recurrent patterns of an expression-trait relationship, predicting the extent to which the trait could be deterministically understood with gene expression data.

genomics

Toward deciphering developmental patterning with deep neural network

Complex biological functions are carried out by the interaction of genes and proteins. Uncovering the gene regulation network behind a function is one of the central themes in biology. Typically, it involves extensive experiments of genetics, biochemistry and molecular biology. In this paper, we show that much of the inference task can be accomplished by a deep neural network (DNN), a form of machine learning or artificial intelligence. Specifically, the DNN learns from the dynamics of the gene expression. The learnt DNN behaves like an accurate simulator of the system, on which one can perform in-silico experiments to reveal the underlying gene network. We demonstrate the method with two examples: biochemical adaptation and the gap-gene patterning in fruit fly embryogenesis. In the first example, the DNN can successfully find the two basic network motifs for adaptation - the negative feedback and the incoherent feed-forward. In the second and much more complex example, the DNN can accurately predict behaviors of essentially all the mutants. Furthermore, the regulation network it uncovers is strikingly similar to the one inferred from experiments. In doing so, we develop methods for deciphering the gene regulation network hidden in the DNN "black box". Our interpretable DNN approach should have broad applications in genotype-phenotype mapping. SignificanceComplex biological functions are carried out by gene regulation networks. The mapping between gene network and function is a central theme in biology. The task usually involves extensive experiments with perturbations to the system (e.g. gene deletion). Here, we demonstrate that machine learning, or deep neural network (DNN), can help reveal the underlying gene regulation for a given function or phenotype with minimal perturbation data. Specifically, after training with wild-type gene expression dynamics data and a few mutant snapshots, the DNN learns to behave like an accurate simulator for the genetic system, which can be used to predict other mutants behaviors. Furthermore, our DNN approach is biochemically interpretable, which helps uncover possible gene regulatory mechanisms underlying the observed phenotypic behaviors.

developmental biology