bioRxiv ScienceSearch

bioRxiv · 10.1101/046474

Probabilistic estimation of short sequence expression using RNA-Seq data and the positional bootstrap

Abstract

When estimating expression of a transcript or part of a transcript using RNA-seq data, it is commonly assumed that reads are generated uniformly from positions within the transcript. While this assumption is acceptable for long transcript sequences where reads from many positions are averaged, it frequently leads to large errors for short sequences, e.g., less than 100 bp. Analysis of short sequences, such as when studying splice junctions and microRNAs, is increasingly important and necessitates addressing errors in short-sequence expression estimation. Indeed, when we examined RNA-seq data from diverse studies, we found that large errors are introduced by variations in RNA-seq coverage due to sequence content, experimental conditions and sample preparation.\n\nWe developed a technique that we call the positional bootstrap, which quantifies the level of uncertainty in expression induced by non-uniform coverage. Unlike methods that attempt to correct for biases in coverage, but do so by making strong assumptions about the form of those biases, the positional bootstrap can quantify the noise induced by all types of bias, including unknown ones. Results obtained using independently generated RNA-seq datasets show that the positional bootstrap increases the accuracy of estimates of alternative splicing levels, tissue-differential alternative splicing and tissue differential expression, by a factor of up to 10.\n\nA Python implementation of the algorithm to quantify splicing levels is freely available from github.com/PSI-Lab/BENTO-Seq.

Explore related subjects

Keep this discovery

BibTeXRIS

Hui Yuan Xiong, Leo J. Lee, Hannes Bretschneider, Jiexin Gao, Nebojsa Jojic, Brendan J. Frey. 2016-04-02. Probabilistic estimation of short sequence expression using RNA-Seq data and the positional bootstrap. https://doi.org/10.1101/046474

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Lysine as a potential low molecular weight angiogen: its clinical, experimental and in-silico validation- A brief study

Globally, the area of angiogenesis is dominated by investigations on anti-angiogenic agents and processes, due to its role in metastatic cancer treatment. Although, the area of ischemic tissue reperfusion is having much bigger demand and foot-mark. Following clinical failure of VEGF (Vascular endothelial growth factor) as a potential agent for induction of a controlled angiogenic response in ischemic tissues and organs, the progress is reasonably quiet as for new low molecular weight (LMW) angiogen molecules and their clinical applications are concerned. Basic amino acid Lysine has been observed to have profound angiogenic property in ischemic tissues, which is controlled, reproducible, time bound and without any accompanying reperfusion damage. In this study, the basic amino acid Lysine has been suggested as a LMW-angiogen, where it has been proposed to have a molecular binding property between VEGF and VEGF receptor (VEGFR). Here, the molecular adhesive hypothesis is being probed and confirmed both in the clinical and lab conditions through induced angiogenic response in tissue repair and in chick chorio allantoic membrane (CAM), respectively; and in dry-docking experiments (in-silico studies).

Molecular Biology

Evolving Notch polyQ tracts reveal possible solenoid interference elements

Polyglutamine (polyQ) tracts in regulatory proteins are extremely polymorphic. As functional elements under selection for length, triplet repeats are prone to DNA replication slippage and indel mutations. Many polyQ tracts are also embedded within intrinsically disordered domains, which are less constrained, fast evolving, and difficult to characterize. To identify structural principles underlying polyQ tracts in disordered regulatory domains, here I analyze deep evolution of metazoan Notch polyQ tracts, which can generate alleles causing developmental and neurogenic defects. I show that Notch features polyQ tract turnover that is restricted to a discrete number of conserved \"polyQ insertion slots\". Notch polyQ insertion slots are: (i) identifiable by an amphipathic \"slot leader\" motif; (ii) conserved as an intact C-terminal array in a 1-to-1 relationship with the N-terminal solenoid-forming ankyrin repeats (ARs); and (iii) enriched in carboxamide residues (Q/N), whose sidechains feature dual hydrogen bond donor and acceptor atoms. Correspondingly, the terminal loop and {beta}-strand of each AR feature conserved carboxamide residues, which would be susceptible to folding interference by hydrogen bonding with residues outside the ARs. I thus suggest that Notch polyQ insertion slots constitute an array of AR interference elements (ARIEs). Notch ARIEs would dynamically compete with the delicate serial folding induced by adjacent ARs. Huntingtin, which harbors solenoid-forming HEAT repeats, also possesses a similar number of polyQ insertion slots. These results strongly suggest that intrinsically disordered interference arrays featuring carboxamide and polyQ enrichment are coupled proteodynamic modulators of solenoids.\n\nSIGNIFICANCENeurodegenerative disorders are often caused by expanded polyglutamine (polyQ) tracts embedded in the disordered regions of regulatory proteins, which are difficult to characterize structurally. To identify functional principles underlying polyQ tracts in disordered regulatory domains, I analyze evolution of the Notch protein, which can generate polyQ-related alleles causing neurodevelopmental defects. I show that Notch evolves polyQ tracts that come and go in a few conserved \"polyQ insertion slots\". Several features suggest these slots are ankyrin repeat (AR) interference elements, which dynamically compete with the delicate solenoid formed by Notch. Huntingtin, whose polyQ expansions causes Huntingtons Disease in humans, also has solenoid-forming modules and polyQ insertion slots, suggesting a common architectural principle underlies solenoid-forming polyQ-rich proteins.

Molecular Biology

Multiplex gene editing by CRISPR-Cpf1 through autonomous processing of a single crRNA array

Microbial CRISPR-Cas defense systems have been adapted as a platform for genome editing applications built around the RNA-guided effector nucleases, such as Cas9. We recently reported the characterization of Cpf1, the effector nuclease of a novel type V-A CRISPR system, and demonstrated that it can be adapted for genome editing in mammalian cells (Zetsche et al., 2015). Unlike Cas9, which utilizes a trans-activating crRNA (tracrRNA) as well as the endogenous RNaseIII for maturation of its dual crRNA:tracrRNA guides (Deltcheva et al., 2011), guide processing of the Cpf1 system proceeds in the absence of tracrRNA or other Cas (CRISPR associated) genes (Zetsche et al., 2015) (Figure 1a), suggesting that Cpf1 is sufficient for pre-crRNA maturation. This has important implications for genome editing, as it would provide a simple route to multiplex targeting. Here, we show for two Cpf1 orthologs that no other factors are required for array processing and demonstrate multiplex gene editing in mammalian cells as well as in the mouse brain by using a designed single CRISPR array.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=164 SRC=\"FIGDIR/small/049122_fig1.gif\" ALT=\"Figure 1\">\nView larger version (35K):\norg.highwire.dtl.DTLVardef@1f0c15corg.highwire.dtl.DTLVardef@126a6fdorg.highwire.dtl.DTLVardef@9d6a2eorg.highwire.dtl.DTLVardef@a61f75_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1C_FLOATNO Cpf1 mediates processing of pre-crRNA. (a) Schematic of pre-crRNA processing for Cas9 and Cpf1. Cleavage sites indicated with red triangle. (b) In vitro processing of FnCpf1 pre-crRNA transcript (80 nM) with purified AsCpf1 or LbCpf1 protein ([~]320 nM). In the presence of Cpf1 nuclease the pre-crRNA was cleaved in a distinct pattern, indicating cleavage at similar sequence motifs. RNA molecules without Cpf1 DR features where not cleaved by Cpf1 (control RNA). (c) RNAseq analysis of FnCpf1 pre-crRNA cleavage products, as shown in (b). A high fraction of sequence reads smaller than 65nt are cleavage products of spacers flanked by DR sequences.\n\nC_FIG

Molecular Biology