bioRxiv ScienceSearch

Biology subjects

Kelly, S.

Publications and source records attributed to Kelly, S..

13 recordsLinked to original sources

Gene duplication accelerates the pace of protein gain and loss from plant organelles.

Introductory paragraphA hallmark of eukaryotic cells is the compartmentalisation of intracellular processes into specialised membrane-bound compartments known as organelles. Plant cells contain several such organelles including the nucleus, chloroplast, mitochondrion, peroxisome, golgi, endoplasmic reticulum and vacuole. Organelle biogenesis and function is dependent on the concerted action of numerous nuclear-encoded proteins which must be imported from the cytosol (or endoplasmic reticulum) where they are made. Using phylogenomic approaches coupled to ancestral state estimation we show that the rate of change in plant organellar proteome content is proportional to the rate of molecular sequence evolution such that the proteomes of chloroplasts and mitochondria lose or gain ~3.2 proteins per million years. We show that these changes in protein targeting have predominantly occurred in genes with regulatory rather than metabolic functions, and thus altered regulatory capacity rather than metabolic function has been the major theme of plant organellar evolution. Finally we show gain and loss of protein targeting occurs at a higher rate following gene duplication events, revealing that gene and genome duplication are a key facilitator of organelle evolution.

evolutionary biology

Defining Inflammatory Cell States in Rheumatoid Arthritis Joint Synovial Tissues by Integrating Single-cell Transcriptomics and Mass Cytometry

To define the cell populations in rheumatoid arthritis (RA) driving joint inflammation, we applied single-cell RNA-seq (scRNA-seq), mass cytometry, bulk RNA-seq, and flow cytometry to sorted T cells, B cells, monocytes, and fibroblasts from 51 synovial tissue RA and osteoarthritis (OA) patient samples. Utilizing an integrated computational strategy based on canonical correlation analysis to 5,452 scRNA-seq profiles, we identified 18 unique cell populations. Combining mass cytometry and transcriptomics together revealed cell states expanded in RA synovia: THY1+HLAhigh sublining fibroblasts (OR=33.8), IL1B+ pro-inflammatory monocytes (OR=7.8), CD11c+T-bet+ autoimmune-associated B cells (OR=5.7), and PD-1+Tph/Tfh (OR=3.0). We also defined CD8+ T cell subsets characterized by GZMK+, GZMB+, and GNLY+ expression. Using bulk and single-cell data, we mapped inflammatory mediators to source cell populations, for example attributing IL6 production to THY1+HLAhigh fibroblasts and naive B cells, and IL1B to pro-inflammatory monocytes. These populations are potentially key mediators of RA pathogenesis.

immunology

High dimensional analyses of cells dissociated from cryopreserved synovial tissue

BackgroundDetailed molecular analyses of cells from rheumatoid arthritis (RA) synovium hold promise in identifying cellular phenotypes that drive tissue pathology and joint damage. The Accelerating Medicines Partnership (AMP) RA/SLE network aims to deconstruct autoimmune pathology by examining cells within target tissues through multiple high-dimensional assays. Robust standardized protocols need to be developed before cellular phenotypes at a single cell level can be effectively compared across patient samples.\n\nMethodsMultiple clinical sites collected cryopreserved synovial tissue fragments from arthroplasty and synovial biopsy in a 10%-DMSO solution. Mechanical and enzymatic dissociation parameters were optimized for viable cell extraction and surface protein preservation for cell sorting and mass cytometry, as well as for reproducibility in RNA sequencing (RNA-seq). Cryopreserved synovial samples were collectively analyzed at a central processing site by a custom-designed and validated 35-marker mass cytometry panel. In parallel, each sample was flow sorted into fibroblast, T cell, B cell, and macrophage suspensions for bulk population RNA-seq and plate-based single cell CEL-Seq2 RNA-seq.\n\nResultsUpon dissociation, cryopreserved synovial tissue fragments yielded a high frequency of viable cells, comparable to samples undergoing immediate processing. Optimization of synovial tissue dissociation across six clinical collection sites with [~]30 arthroplasty and [~]20 biopsy samples yielded a consensus digestion protocol using 100{micro}g/mL of Liberase TL enzyme. This protocol yielded immune and stromal cell lineages with preserved surface markers and minimized variability across replicate RNA-seq transcriptomes. Mass cytometry analysis of cells from cryopreserved synovium distinguished: 1) diverse fibroblast phenotypes, 2) distinct populations of memory B cells and antibody-secreting cells, and 3) multiple CD4+ and CD8+ T cell activation states. Bulk RNA sequencing of sorted cell populations demonstrated robust separation of synovial lymphocytes, fibroblasts, and macrophages. Single cell RNA-seq produced transcriptomes of over 1000 genes/cell, including transcripts encoding characteristic lineage markers identified.\n\nConclusionWe have established a robust protocol to acquire viable cells from cryopreserved synovial tissue with intact transcriptomes and cell surface phenotypes. A centralized pipeline to generate multiple high-dimensional analyses of synovial tissue samples collected across a collaborative network was developed. Integrated analysis of such datasets from large patient cohorts may help define molecular heterogeneity within RA pathology and identify new therapeutic targets and biomarkers.

immunology

STAG: Species Tree Inference from All Genes

Species tree inference is fundamental to our understanding of the evolution of life on earth. However, species tree inference from molecular sequence data is complicated by gene duplication events that limit the availably of suitable data for phylogenetic reconstruction. Here we propose a novel method for species tree inference called STAG that is specifically designed to leverage data from multi-copy gene families. By application to 12 real species datasets sampled from across the eukaryotic domain we demonstrate that species trees inferred from multi-copy gene families are comparable in accuracy to species trees inferred from single-copy orthologues. We further show that the ability to utilise data from multi-copy gene families increases the amount of data available for species tree inference by an average of 8 fold. We reveal that on real species datasets STAG has higher accuracy than other leading methods for species tree inference; including concatenated alignments of protein sequences, ASTRAL & NJst. Finally we show that STAG is fast, memory efficient and scalable and thus suitable for analysis of large multispecies datasets.

evolutionary biology

Clust: automatic extraction of optimal co-expressed gene clusters from gene expression data

Identification of co-expressed gene clusters can provide evidence for genetic or physical interactions between genes. Thus, co-expression clustering is a routine step in large-scale analyses of gene expression data. We show that commonly used clustering methods produce results that substantially disagree with each other, and do not match the biological expectations of co-expressed gene clusters. Furthermore, these clusters can contain up to 50% unreliably assigned genes. Consequently, downstream analyses of these clusters (e.g. functional term enrichment analysis) suffer from high error rates. We present clust, an automated method that solves these problems by extracting clusters that match the biological expectations of co-expressed genes. Using 100 datasets from five model organisms we demonstrate that clusters generated by clust are better than those produced by other methods, both numerically and for use in functional analysis. Finally, we show that clust can simultaneously cluster multiple datasets, enabling users to leverage the large quantity of public expression data for novel comparative analysis.

bioinformatics

OMGene: Mutual improvement of gene models through optimisation of evolutionary conservation

BackgroundThe accurate determination of the genomic coordinates for a given gene - its gene model - is of vital importance to the utility of its annotation, and the accuracy of bioinformatic analyses derived from it. Currently-available methods of computational gene prediction, while on the whole successful, often disagree on the model for a given predicted gene, with some or all of the variant gene models failing to match the biologically observed structure. Many prediction methods can be bolstered by using experimental data such as RNA-seq and mass spectrometry. However, these resources are not always available, and rarely give a comprehensive portrait of an organisms transcriptome due to temporal and tissue-specific expression profiles.\n\nResultsOrthology between genes provides evolutionary evidence to guide the construction of gene models. OMGene (Optimise My Gene) aims to optimise gene models in the absence of experimental data by optimising the derived amino acid alignments for gene models within orthogroups. Using RNA-seq data sets from plants and fungi, considering intron/exon junction representation and exon coverage, and assessing the intra-orthogroup consistency of subcellular localisation predictions, we demonstrate the utility of OMGene for improving gene models in annotated genomes.\n\nConclusionsWe show that significant improvements in the accuracy of gene model annotations can be made in both established and de novo annotated genomes by leveraging information from multiple species.

genomics

The amount of nitrogen used for photosynthesis governs molecular evolution in plants

BackgroundGenome and transcript sequences are composed of long strings of nucleotide monomers (A, C, G and T/U) that require different quantities of nitrogen atoms for biosynthesis.\n\nResultsHere it is shown that the strength of selection acting on transcript nitrogen content is determined by the amount of nitrogen plants require to conduct photosynthesis. Specifically, plants that require more nitrogen to conduct photosynthesis experience stronger selection on transcript sequences to use synonymous codons that cost less nitrogen to biosynthesise. It is further shown that the strength of selection acting on transcript nitrogen cost constrains molecular sequence evolution such that genes experiencing stronger selection evolve at a slower rate.\n\nConclusionsTogether these findings reveal that the plant molecular clock is set by photosynthetic efficiency, and provide a mechanistic explanation for changes in plant speciation rates that occur concomitant with improvements in photosynthetic efficiency and changes in the environment such as light, temperature, and atmospheric CO2 concentration.

evolutionary biology

Wide sampling of natural diversity identifies novel molecular signatures of C4 photosynthesis

Introductory paragraphMuch of biology is associated with convergent traits, and it is challenging to determine the extent to which underlying molecular mechanisms are shared across phylogeny. By analyzing plants representing eighteen independent origins of C4 photosynthesis, we quantified the extent to which this convergent trait utilises identical molecular mechanisms. We demonstrate that biochemical changes that characterise C4 species are recovered by this process, and expand the paradigm by four metabolic pathways not previously associated with C4 photosynthesis. Furthermore, we show that expression of many genes that distinguish C3 and C4 species respond to low CO2, providing molecular evidence that reduction in atmospheric CO2 was a driver for C4 evolution. Thus the origin and architecture of complex traits can be derived from transcriptome comparisons across natural diversity.

plant biology

STRIDE: Species Tree Root Inference From Gene Duplication Events

The correct interpretation of a phylogenetic tree is dependent on it being correctly rooted. A gene duplication event at the base of a clade of species is synapamorphic, and thus excludes the root of the species tree from that clade. We present STRIDE, a fast, effective, and outgroup-free method for species tree root inference from gene duplication events. STRIDE identifies sets of well-supported gene duplication events from cohorts of gene trees, and analyses these events to infer a probability distribution over an unrooted species tree for the location of the true root. We show that STRIDE infers the correct root of the species tree for a large range of simulated and real species sets. We demonstrate that the novel probability model implemented in STRIDE can accurately represent the ambiguity in species tree root assignment for datasets where information is limited. Furthermore, application of STRIDE to inference of the origin of the eukaryotic tree resulted in a root probability distribution that was consistent with, but unable to distinguish between, leading hypotheses for the origin of the eukaryotes. In summary, STRIDE is a fast, scalable, and effective method for species tree root inference from genome scale data.

evolutionary biology

Organ-Specific NLR Resistance Gene Expression Varies With Plant Symbiotic Status

Nucleotide-binding site leucine-rich repeat resistance genes (NLRs) allow plants to detect microbial effectors. We hypothesized that NLR expression patterns would reflect organ-specific differences in effector challenge and tested this by carrying out a meta-analysis of expression data for 1,235 NLRs from 9 plant species. We found stable NLR root/shoot expression ratios within species, suggesting organ-specific hardwiring of NLR expression patterns in anticipation of distinct challenges. Most monocot and dicot plant species preferentially expressed NLRs in roots. In contrast, Brassicaceae species, including oilseed rape and the model plant Arabidopsis thaliana, were unique in showing NLR expression skewed towards the shoot across multiple phylogenetically distinct groups of NLRs. The Brassicaceae NLR expression shift coincides with loss of the endomycorrhization pathway, which enables intracellular root infection by symbionts. We propose that its loss offer two likely explanations for the unusual Brassicaceae NLR expression pattern: loss of NLR-guarded symbiotic components and elimination of constraints on general root defences associated with exempting symbionts from targeting. This hypothesis is consistent with the existence of Brassicaceae-specific receptors for conserved microbial molecules and suggests that Brassicaceae species are rich sources of unique antimicrobial root defences.

plant biology

Selection-Driven Cost-Efficiency Optimisation Of Transcript Sequences Determines The Rate Of Gene Sequence Evolution In Bacteria

BackgroundMost amino acids are encoded by multiple synonymous codons. However synonymous codons are not used equally and this biased codon use varies between different organisms. It has previously been shown that both selection acting to increase codon translational efficiency and selection acting to decrease codon biosynthetic cost contribute to differences in codon bias. However, it is unknown how these two factors interact or how they affect molecular sequence evolution.\n\nResultsThrough analysis of 1,320 bacterial genomes we show that bacterial genes are subject to multi-objective selection-driven optimisation of codon use. Here, selection acts to simultaneously decrease transcript biosynthetic cost and increase transcript translational efficiency, with highly expressed genes under the greatest selection. This optimisation is not simply a consequence of the more translationally efficient codons being less expensive to synthesise. Instead, we show that tRNA gene copy number alters the cost-efficiency trade-off of synonymous codons such that for many species such that selection acting on transcript biosynthetic cost and translational efficiency act in opposition. Finally, we show that genes highly optimised to reduce cost and increase efficiency show reduced rates of synonymous and non-synonymous mutation.\n\nConclusionsThis analysis provides a simple mechanistic explanation for variation in evolutionary rate between genes that depends on selection-driven cost-efficiency optimisation of the transcript. These findings reveal how optimisation of resource allocation to mRNA synthesis is a critical factor that determines both the evolution and composition of genes.

evolutionary biology

OrthoFiller: utilising data from multiple species to improve the completeness of genome annotations.

BackroundComplete and accurate annotation of sequenced genomes is of paramount importance to their utility and analysis. Differences in gene prediction pipelines mean that genome sequences for a species can differ considerably in the quality and quantity of their predicted genes. Furthermore, genes that are present in genome sequences sometimes fail to be detected by computational gene prediction methods. Erroneously unannotated genes can lead to oversights and inaccurate assertions in biological investigations, especially for smaller-scale genome projects which rely heavily on computational prediction.\n\nResultsHere we present OrthoFiller, a tool designed to address the problem of finding and adding such missing genes to genome annotations. OrthoFiller leverages information from multiple related species to identify those genes whose existence can be verified through comparison with known gene families, but which have not been predicted. By simulating missing gene annotations in real sequence datasets from both plants and fungi we demonstrate the accuracy and utility of OrthoFiller for finding missing genes and improving genome annotation. Furthermore, we show that applying OrthoFiller to existing \"complete\" genome annotations can identify and correct substantial numbers of erroneously missing genes in these two sets of species.\n\nConclusionsWe show that significant improvements in the completeness of genome annotations can be made by leveraging information from multiple species.

bioinformatics

HiC-bench: comprehensive and reproducible Hi-C data analysis designed for parameter exploration and benchmarking

BackgroundChromatin conformation capture techniques have evolved rapidly over the last few years and have provided new insights into genome organization at an unprecedented resolution. Analysis of Hi-C data is complex and computationally intensive involving multiple tasks and requiring robust quality assessment. This has led to the development of several tools and methods for processing Hi-C data. However, most of the existing tools do not cover all aspects of the analysis and only offer few quality assessment options. Additionally, availability of a multitude of tools makes scientists wonder how these tools and associated parameters can be optimally used, and how potential discrepancies can be interpreted and resolved. Most importantly, investigators need to be ensured that slight changes in parameters and/or methods do not affect the conclusions of their studies.\n\nResultsTo address these issues (compare, explore and reproduce), we introduce HiC-bench, a configurable computational platform for comprehensive and reproducible analysis of Hi-C sequencing data. HiC-bench performs all common Hi-C analysis tasks, such as alignment, filtering, contact matrix generation and normalization, identification of topological domains, scoring and annotation of specific interactions using both published tools and our own. We have also embedded various tasks that perform quality assessment and visualization. HiC-bench is implemented as a data flow platform with an emphasis on analysis reproducibility. Additionally, the user can readily perform parameter exploration and comparison of different tools in a combinatorial manner that takes into account all desired parameter settings in each pipeline task. This unique feature facilitates the design and execution of complex benchmark studies that may involve combinations of multiple tool/parameter choices in each step of the analysis. To demonstrate the usefulness of our platform, we performed a comprehensive benchmark of existing and new TAD callers exploring different matrix correction methods, parameter settings and sequencing depths. Users can extend our pipeline by adding more tools as they become available.\n\nConclusionsHiC-bench consists an easy-to-use and extensible platform for comprehensive analysis of Hi-C datasets. We expect that it will facilitate current analyses and help scientists formulate and test new hypotheses in the field of three-dimensional genome organization.

bioinformatics