bioRxiv ScienceSearch

Biology subjects

John Marioni

Publications and source records attributed to John Marioni.

3 recordsLinked to original sources

Structure and evolutionary history of a large family of NLR proteins in the zebrafish

Animals and plants have evolved a range of mechanisms for recognizing noxious substances and organisms. A particular challenge, most successfully met by the adaptive immune system in vertebrates, is the specific recognition of potential pathogens, which themselves evolve to escape recognition. A variety of genomic and evolutionary mechanisms shape large families of proteins dedicated to detecting pathogens and create the diversity of binding sites needed for epitope recognition. One family involved in innate immunity are the NACHT-domain-and Leucine-Rich-Repeat-containing (NLR) proteins. Mammals have a small number of NLR proteins, which are involved in first-line immune defense and recognize several conserved molecular patterns. However, there is no evidence that they cover a wider spectrum of differential pathogenic epitopes. In other species, mostly those without adaptive immune systems, NLRs have expanded into very large families. A family of nearly 400 NLR proteins is encoded in the zebrafish genome. They are subdivided into four groups defined by their NACHT and effector domains, with a characteristic overall structure that arose in fishes from a fusion of the NLR domains with a domain used for immune recognition, the B30.2 domain. The majority of the genes are located on one chromosome arm, interspersed with other large multi-gene families, including a new family encoding proteins with multiple tandem arrays of Zinc fingers. This chromosome arm may be a hot spot for evolutionary change in the zebrafish genome. NLR genes not on this chromosome tend to be located near chromosomal ends.\n\nExtensive duplication, loss of genes and domains, exon shuffling and gene conversion acting differentially on the NACHT and B30.2 domains have shaped the family. Its four groups, which are conserved across the fishes, are homogenised within each group by gene conversion, while the B30.2 domain is subject to gene conversion across the groups. Evidence of positive selection on diversifying mutations in the B30.2 domain, probably driven by pathogen interactions, indicates that this domain rather than the LRRs acts as a recognition domain. The NLR-B30.2 proteins represent a new family with diversity in the specific recognition module that is present in fishes in spite of the parallel existence of an adaptive immune system.

Genomics

Structure and evolutionary history of a large family of NLR proteins in the zebrafish

NACHT- and Leucine-Rich-Repeat-containing domain (NLR) proteins act as cytoplasmic sensors for pathogen- and danger-associated molecular patterns and are found throughout the plant and animal kingdoms. In addition to having a small set of conserved NLRs, the genomes in some animal lineages contain massive expansions of this gene family. One of these arose in fishes, after the creation of a gene fusion that combined the core NLR domains with another domain used for immune recognition, the PRY/SPRY or B30.2 domain. We have analysed the expanded NLR gene family in zebrafish, which contains 368 genes, and studied its evolutionary history. The encoded proteins share a defining overall structure, but individual domains show different evolutionary trajectories. Our results suggest gene conversion homogenizes NACHT and B30.2 domain sequences among different gene subfamilies, however, the functional implications of its action remains unclear. The majority of the genes are located on the long arm of chromosome 4, interspersed with several other large multi-gene families, including a new family encoding proteins with multiple tandem arrays of Zinc fingers. This suggests that chromosome 4 may be a hotspot for rapid evolutionary change in zebrafish.

Evolutionary Biology

iRAP - an integrated RNA-seq Analysis Pipeline

RNA-sequencing (RNA-Seq) has become the technology of choice for whole-transcriptome profiling. However, processing the millions of sequence reads generated requires considerable bioinformatics skills and computational resources. At each step of the processing pipeline many tools are available, each with specific advantages and disadvantages. While using a specific combination of tools might be desirable, integrating the different tools can be time consuming, often due to specificities in the formats of input/output files required by the different programs. Here we present iRAP, an integrated RNA-seq analysis pipeline that allows the user to select and apply their preferred combination of existing tools for mapping reads, quantifying expression, testing for differential expression. iRAP also includes multiple tools for gene set enrichment analysis and generates web browsable reports of the results obtained in the different stages of the pipeline. Depending upon the application, iRAP can be used to quantify expression at the gene, exon or transcript level. iRAP is aimed at a broad group of users with basic bioinformatics training and requires little experience with the command line. Despite this, it also provides more advanced users with the ability to customise the options used by their chosen tools.

Bioinformatics