bioRxiv ScienceSearch

Biology subjects

Roderic Guigo

Publications and source records attributed to Roderic Guigo.

7 recordsLinked to original sources

ChimPipe: Accurate detection of fusion genes and transcription-induced chimeras from RNA-seq data

BackgroundChimeric transcripts are commonly defined as transcripts linking two or more different genes in the genome, and can be explained by various biological mechanisms such as genomic rearrangement, read-through or trans-splicing, but also by technical or biological artefacts. Several studies have shown their importance in cancer, cell pluripotency and motility. Many programs have recently been developed to identify chimeras from Illumina RNA-seq data (mostly fusion genes in cancer). However outputs of different programs on the same dataset can be widely inconsistent, and tend to include many false positives. Other issues relate to simulated datasets restricted to fusion genes, real datasets with limited numbers of validated cases, result inconsistencies between simulated and real datasets, and gene rather than junction level assessment.\n\nResultsHere we present ChimPipe, a modular and easy-to-use method to reliably identify chimeras from paired-end Illumina RNA-seq data. We have also produced realistic simulated datasets for three different read lengths, and enhanced two gold-standard cancer datasets by associating exact junction points to validated gene fusions. Benchmarking ChimPipe together with four other state-of-the-art tools on this data showed ChimPipe to be the top program at identifying exact junction coordinates for both kinds of datasets, and the one showing the best trade-off between sensitivity and precision. Applied to 106 ENCODE human RNA-seq datasets, ChimPipe identified 137 high confidence chimeras connecting the protein coding sequence of their parent genes. In subsequent experiments, three out of four predicted chimeras, two of which recurrently expressed in a large majority of the samples, could be validated. Cloning and sequencing of the three cases revealed several new chimeric transcript structures, 3 of which with the potential to encode a chimeric protein for which we hypothesized a new role.\n\nConclusionsChimPipe combines spanning and paired end RNA-seq reads to detect any kind of chimeras, including read-throughs, and shows an excellent trade-off between sensitivity and precision. The chimeras found by ChimPipe can be validated in-vitro with high accuracy.

Bioinformatics

Discovery of Cancer Driver Long Noncoding RNAs across 1112 Tumour Genomes: New Candidates and Distinguishing Features.

Long noncoding RNAs (lncRNAs) represent a vast unexplored genetic space that may hold missing drivers of tumourigenesis, but few such \"driver lncRNAs\" are known. Until now, they have been discovered through changes in expression, leading to problems in distinguishing between causative roles and passenger effects. We here present a different approach for driver lncRNA discovery using mutational patterns in tumour DNA. Our pipeline, ExInAtor, identifies genes with excess load of somatic single nucleotide variants (SNVs) across panels of tumour genomes. Heterogeneity in mutational signatures between cancer types and individuals is accounted for using a simple local trinucleotide background model, which yields high precision and low computational demands. We use ExInAtor to predict drivers from the GENCODE annotation across 1112 entire genomes from 23 cancer types. Using a stratified approach, we identify 15 high-confidence candidates: 9 novel and 6 known cancer-related genes, including MALAT1, NEAT1 and SAMMSON. Both known and novel driver lncRNAs are distinguished by elevated gene length, evolutionary conservation and expression. We have presented a first catalogue of mutated lncRNA genes driving cancer, which will grow and improve with the application of ExInAtor to future tumour genome projects.

Cancer Biology

Scalable Design of Paired CRISPR Guide RNAs for Genomic Deletion

Using CRISPR/Cas9, diverse genomic elements may be studied in their endogenous context. Pairs of single guide RNAs (sgRNAs) are used to delete regulatory elements and small RNA genes, while longer RNAs can be silenced through promoter deletion. We here present CRISPETa, a bioinformatic pipeline for flexible and scalable paired sgRNA design based on an empirical scoring model. Multiple sgRNA pairs are returned for each target. Any number of targets can be analyzed in parallel, making CRISPETa equally appropriate for studies of individual elements, or complex library screens. Fast run-times are achieved using a precomputed off-target database. sgRNA pair designs are output in a convenient format for visualisation and oligonucleotide ordering. We present a series of pre-designed, high-coverage library designs for entire classes of non-coding elements in human, mouse, zebrafish, Drosophila and C. elegans. Using an improved version of the DECKO deletion vector, together with a quantitative deletion assay, we test CRISPETa designs by deleting an enhancer and exonic fragment of the MALAT1 oncogene. These achieve efficiencies of [≥]50%, resulting in production of mutant RNA. CRISPETa will be useful for researchers seeking to harness CRISPR for targeted genomic deletion, in a variety of model organisms, from single-target to high-throughput scales.

Genomics

The discovery potential of RNA processing profiles

Small non-coding RNAs are highly abundant molecules that regulate essential cellular processes and are classified according to sequence and structure. Here we argue that read profiles from size-selected RNA sequencing capture the post-transcriptional processing specific to each RNA family, thereby providing functional information independently of sequence and structure. We developed SeRPeNT, the first unsupervised computational method that exploits reproducibility across replicates and uses dynamic time-warping and density-based clustering algorithms to identify, characterize and compare small non-coding RNAs (sncRNAs) by harnessing the power of read profiles. We applied SeRPeNT to: a) generate an extended human annotation with 671 new sncRNAs from known classes and 131 from new potential classes, b) show pervasive differential processing between cell compartments and c) predict new molecules with miRNA-like behaviour from snoRNA, tRNA and long non-coding RNA precursors, potentially dependent on the miRNA biogenesis pathway. Furthermore, we validated experimentally four predicted novel non-coding RNAs: a miRNA, a snoRNA-derived miRNA, a processed tRNA and a new uncharacterized sncRNA. SeRPeNT facilitates fast and accurate discovery and characterization of small non-coding RNAs at unprecedented scale. SeRPeNT code is available under the MIT license at https://github.com/comprna/SeRPeNT.

Genomics

Evolution of selenophosphate synthetases: emergence and relocation of function through independent duplications and recurrent subfunctionalization

SPS catalyzes the synthesis of selenophosphate, the selenium donor for the synthesis of the amino acid selenocysteine (Sec), incorporated in selenoproteins in response to the UGA codon. SPS is unique among proteins of the selenoprotein biosynthesis machinery in that it is, in many species, a selenoprotein itself, although, as in all selenoproteins, Sec is often replaced by cysteine (Cys). In metazoan genomes we found, however, SPS genes with lineage specific substitutions other than Sec or Cys. Our results show that these non-Sec, non-Cys SPS genes originated through a number of independent gene duplications of diverse molecular origin from an ancestral selenoprotein SPS gene. Although of independent origin, complementation assays in fly mutants show that these genes share a common function, which most likely emerged in the ancestral metazoan gene. This function appears to be unrelated to selenophosphate synthesis, since all genomes encoding selenoproteins contain Sec or Cys SPS genes (SPS2), but those containing only non-Sec, non-Cys SPS genes (SPS1) do not encode selenoproteins. Thus, in SPS genes, through parallel duplications and subsequent convergent subfunctionalization, two functions initially carried by a single gene are recurrently segregated at two different loci. RNA structures enhancing the readthrough of the Sec-UGA codon in SPS genes, which may be traced back to prokaryotes, played a key role in this process. The SPS evolutionary history in metazoans constitute a remarkable example of the emergence and evolution of gene function. We have been able to trace this history with unusual detail thanks to the singular feature of SPS genes, wherein the amino acid at a single site determines protein function, and, ultimately, the evolutionary fate of an entire class of genes.

Genomics

Widespread localisation of long noncoding RNAs to ribosomes: Distinguishing features and evidence for regulatory roles.

The function of long noncoding RNAs (lncRNAs) depends on their location within the cell. While most studies to date have concentrated on their nuclear roles in transcriptional regulation, evidence is mounting that lncRNA also have cytoplasmic roles. Here we comprehensively map the cytoplasmic and ribosomal lncRNA population in a human cell. Three-quarters (74%) of lncRNAs are detected in the cytoplasm, the majority of which (62%) preferentially cofractionate with polyribosomes. Ribosomal lncRNA are highly expressed across tissues, under purifying evolutionary selection, and have cytoplasmic-to-nuclear ratios comparable to mRNAs and consistent across cell types. LncRNAs may be classified into three groups by their ribosomal interaction: non-ribosomal cytoplasmic lncRNAs, and those associated with either heavy or light polysomes. A number of mRNA-like features destin lncRNA for light polysomes, including capping and 5UTR length, but not cryptic open reading frames or polyadenylation. Surprisingly, exonic retroviral sequences antagonise recruitment. In contrast, it appears that lncRNAs are recruited to heavy polysomes through basepairing to mRNAs. Finally, we show that the translation machinery actively degrades lncRNA. We propose that light polysomal lncRNAs are translationally engaged, while heavy polysomal lncRNAs are recruited indirectly. These findings point to extensive and reciprocal regulatory interactions between lncRNA and the translation machinery.

Genomics

Enhanced Transcriptome Maps from Multiple Mouse Tissues Reveal Evolutionary Constraint in Gene Expression for Thousands of Genes

We characterized by RNA-seq the transcriptional profiles of a large and heterogeneous collection of mouse tissues, augmenting the mouse transcriptome with thousands of novel transcript candidates. Comparison with transcriptome profiles obtained in human cell lines reveals substantial conservation of transcriptional programs, and uncovers a distinct class of genes with levels of expression across cell types and species, that have been constrained early in vertebrate evolution. This core set of genes capture a substantial and constant fraction of the transcriptional output of mammalian cells, and participates in basic functional and structural housekeeping processes common to all cell types. Perturbation of these constrained genes is associated with significant phenotypes including embryonic lethality and cancer. Evolutionary constraint in gene expression levels is not reflected in the conservation of the genomic sequences, but is associated with strong and conserved epigenetic marking, as well as to a characteristic post-transcriptional regulatory program in which sub-cellular localization and alternative splicing play comparatively large roles.

Bioinformatics