bioRxiv ScienceSearch

Biology subjects

Haerty, W.

Publications and source records attributed to Haerty, W..

7 recordsLinked to original sources

Identification of functional long non-coding RNAs in C. elegans

BackgroundFunctional characterisation of the compact genome of the model organism Caenorhabditis elegans remains incomplete despite its sequencing twenty years ago. The last decade of research has seen a tremendous increase in the number of non-coding RNAs identified in various organisms. While we have mechanistic understandings of small non-coding RNA pathways, long non-coding RNAs represent a diverse class of active transcripts whose function remains less well characterised.\n\nResultsBy analysing hundreds of published transcriptome datasets, we annotated 3,397 potential lncRNAs including 146 multi-exonic loci that showed increased nucleotide conservation and GC content relative to other non-coding regions. Using CRISPR / Cas9 genome editing we generated deletion mutants for ten long non-coding RNA loci. Using automated microscopy for in-depth phenotyping, we show that six of the long non-coding RNA loci are required for normal development and fertility. Using RNA interference mediated gene knock-down, we provide evidence that for two of the long non-coding RNA loci, the observed phenotypes are dependent on the corresponding RNA transcripts.\n\nConclusionsOur results highlight that a large section of the non-coding regions of the C. elegans genome remain unexplored. Based on our in vivo analysis of a selection of high-confidence lncRNA loci, we expect that a significant proportion of these high-confidence regions is likely to have biological function at either the genomic or the transcript level.

genetics

Cerox1 and microRNA-488-3p noncoding RNAs jointly regulate mitochondrial complex I catalytic activity

To generate energy efficiently, the cell is uniquely challenged to co-ordinate the abundance of electron transport chain protein subunits expressed from both nuclear and mitochondrial genomes. How an effective stoichiometry of this many constituent subunits is co-ordinated post-transcriptionally remains poorly understood. Here we show that Cerox1, an unusually abundant cytoplasmic long noncoding RNA (lncRNA), modulates the levels of mitochondrial complex I subunit transcripts in a manner that requires binding to microRNA-488-3p. Increased abundance of Cerox1 cooperatively elevates complex I subunit protein abundance and enzymatic activity, decreases reactive oxygen species production, and protects against the complex I inhibitor rotenone. Cerox1 function is conserved across placental mammals: human and mouse orthologues effectively modulate complex I enzymatic activity in mouse and human cells, respectively. Cerox1 is the first lncRNA demonstrated, to our knowledge, to regulate mitochondrial oxidative phosphorylation (OXPHOS) and, with miR-488-3p, represent novel targets for the modulation of complex I activity.

biochemistry

Long-read sequencing reveals the splicing profile of the calcium channel gene CACNA1C in human brain

RNA splicing is a key mechanism linking genetic variation with psychiatric disorders. Splicing profiles are particularly diverse in brain and difficult to accurately identify and quantify. We developed a new approach to address this challenge, combining long-range PCR and nanopore sequencing with a novel bioinformatics pipeline. We identify the full-length coding transcripts of CACNA1C in human brain. CACNA1C is a psychiatric risk gene that encodes the voltage-gated calcium channel CaV1.2. We show that CACNA1Cs transcript profile is substantially more complex than appreciated, identifying 38 novel exons and 241 novel transcripts. Importantly, many of the novel variants are abundant, and predicted to encode channels with altered function. The splicing profile varies between brain regions, especially in cerebellum. We demonstrate that human transcript diversity (and thereby protein isoform diversity) remains under-characterised, and provide a feasible and cost-effective methodology to address this. A detailed understanding of isoform diversity will be essential for the translation of psychiatric genomic findings into pathophysiological insights and novel psychopharmacological targets.

neuroscience

The evolutionary dynamics of microRNAs in domestic mammals

MicroRNAs are crucial regulators of gene expression found across both the plant and animal kingdoms. While the numberof annotated microRNAs deposited in miRBase has greatly increased in recent years, few studies provided comparative analyses across sets of related species, or investigated the role of microRNAs in the evolution of gene regulation.\n\nWe generated small RNA libraries across 5 mammalian species (cow, dog, horse, pig and rabbit) from 4 different tissues (brain, heart, kidney and testis). We identified 1675 miRBase and 413 novel microRNAs by manually curating the set of computational predictions obtained from miRCat and miRDeep2.\n\nOur dataset spanning five species has enabled us to investigate the molecular mechanisms and selective pressures driving the evolution of microRNAs in mammals. We highlight the important contributions of intronic sequences (366 orthogroups), duplication events (135 orthogroups) and repetitive elements (37 orthogroups) in the emergence of new microRNA loci.\n\nWe use this framework to estimate the patterns of gains and losses across the phylogeny, and observe high levels of microRNA turnover. Additionally, the identification of lineage-specific losses enables the characterisation of the selective constraints acting on the associated target sites.\n\nCompared to the miRBase subset, novel microRNAs tend to be more tissue specific. 20 percent of novel orthogroups are restricted to the brain, and their target repertoires appear to be enriched for neuron activity and differentiation processes. These findings may reflect an important role for young microRNAs in the evolution of brain expression plasticity.\n\nMany seed sequences appear to be specific to either the cow or the dog. Analyses on the associated targets highlightthe presence of several genes under artificial positive selection, suggesting an involvement of these microRNAs in the domestication process.\n\nAltogether, we provide an overview on the evolutionary mechanisms responsible for microRNA turnover in 5 domestic species, and their possible contribution to the evolution of gene regulation.

evolutionary biology

A quantitative model for characterizing the evolutionary history of mammalian gene expression

Characterizing the evolutionary history of a genes expression profile is a critical component for understanding the relationship between genotype, expression, and phenotype. However, it is not well-established how best to distinguish the different evolutionary forces acting on gene expression. Here, we use RNA-seq across 7 tissues from 17 mammalian species to show that expression evolution across mammals is accurately modeled by the Ornstein-Uhlenbeck (OU) process. This stochastic process models expression trajectories across time as Gaussian distributions whose variance is parameterized by the rate of genetic drift and strength of stabilizing selection. We use these mathematical properties to identify expression pathways under neutral, stabilizing, and directional selection, and quantify the extent of selective pressure on a genes expression. We further detect deleterious expression levels outside expected evolutionary distributions in expression data from individual patients. Our work provides a statistical framework for interpreting expression data across species and in disease.\n\nOne Sentence SummaryWe demonstrate the power of a stochastic model for quantifying selective pressure on expression and estimating evolutionary distributions of optimal gene expression.

genomics

Immune receptors with exogenous domain fusions form evolutionary hotspots in grass genomes

BackgroundThe plant immune system is innate, encoded in the germline. Using it efficiently, plants are capable of recognizing a diverse range of rapidly evolving pathogens. A recently described phenomenon shows that plant immune receptors are able to recognize pathogen effectors through the acquisition of exogenous protein domains from other plant genes.\n\nResultsWe showed that plant immune receptors with integrated domains are distributed unevenly across their phylogeny in grasses. Using phylogenetic analysis, we uncovered a major integration clade, whose members underwent repeated independent integration events producing diverse fusions. This clade is ancestral in grasses with members often found on syntenic chromosomes. Analyses of these fusion events revealed that homologous receptors can be fused to diverse domains. Furthermore, we discovered a 43 amino acids long motif that was associated with this dominant integration clade and was located immediately upstream of the fusion site. Sequence analysis revealed that DNA transposition and/or ectopic recombination are the most likely mechanisms of NLR-ID formation.\n\nConclusionsThe identification of this subclass of plant immune receptors that is naturally adapted to new domain integration will inform biotechnological approaches for generating synthetic receptors with novel pathogen baits.

evolutionary biology

GeneSeqToFamily: the Ensembl Compara GeneTrees pipeline as a Galaxy workflow

BackgroundGene duplication is a major factor contributing to evolutionary novelty, and the contraction or expansion of gene families has often been associated with morphological, physiological and environmental adaptations. The study of homologous genes helps us to understand the evolution of gene families. It plays a vital role in finding ancestral gene duplication events as well as identifying genes that have diverged from a common ancestor under positive selection. There are various tools available, such as MSOAR, OrthoMCL and HomoloGene, to identify gene families and visualise syntenic information between species, providing an overview of syntenic regions evolution at the family level. Unfortunately, none of them provide information about structural changes within genes, such as the conservation of ancestral exon boundaries amongst multiple genomes. The Ensembl GeneTrees computational pipeline generates gene trees based on coding sequences and provides details about exon conservation, and is used in the Ensembl Compara project to discover gene families.\n\nFindingsA certain amount of expertise is required to configure and run the Ensembl Compara GeneTrees pipeline via command line. Therefore, we have converted the command line Ensembl Compara GeneTrees pipeline into a Galaxy workflow, called GeneSeqToFamily, and provided additional functionality. This workflow uses existing tools from the Galaxy ToolShed, as well as providing additional wrappers and tools that are required to run the workflow.\n\nConclusionsGeneSeqToFamily represents the Ensembl Compara pipeline as a set of interconnected Galaxy tools, so they can be run interactively within the Galaxys user-friendly workflow environment while still providing the flexibility to tailor the analysis by changing configurations and tools if necessary. Additional tools allow users to subsequently visualise the gene families produced by the workflow, using the Aequatus.js interactive tool, which has been developed as part of the Aequatus software project.

bioinformatics