bioRxiv ScienceSearch

Biology subjects

Flicek, P.

Publications and source records attributed to Flicek, P..

12 recordsLinked to original sources

Adaptation of proteins to the cold in Antarctic fish: A role for Methionine?

The evolution of antifreeze glycoproteins has enabled notothenioid fish to flourish in the freezing waters of the Southern Ocean. Whilst successful at the biodiversity level to life in the cold, paradoxically at the cellular level these stenothermal animals have problems producing, folding and degrading proteins at their ambient temperatures of down to -1.86{degrees}C. In this first multi-species transcriptome comparison of the amino acid composition of notothenioid proteins with temperate teleost proteins, we show that, unlike psychrophilic bacteria, Antarctic fish provide little evidence for the mass alteration of protein amino acid composition to enhance protein folding and reduce protein denaturation in the cold. The exception was the significant over-representation of positions where leucine in temperate fish proteins was replaced by methionine in the notothenioid orthologues. Although methionine may increase stability in critical proteins, we hypothesise that a more likely explanation for the extra methionines is that they have been preferentially assimilated into the genome because they act as redox sensors. This redox hypothesis is supported by the enrichment of duplicated genes within the notothenioid transcriptomes which centre around Mapk signalling, a major pathway in the cellular cascades associated with responses to environmental stress. Whilst notothenioid fish show cold-associated problems with protein homeostasis, they may have modified only a selected number of biochemical pathways to work efficiently below 0{degrees}C. Even a slight warming of the Southern Ocean might disrupt the critical functions of this handful of key pathways with considerable impacts for the functioning of this ecosystem in the future.

evolutionary biology

Pseudogenes in the mouse lineage: transcriptional activity and strain-specific history

Pseudogenes are ideal markers of genome remodeling. In turn, the mouse is an ideal platform for studying them, particularly with the availability of developmental transcriptional data and the sequencing of 18 strains. Here, we present a comprehensive genome-wide annotation of the pseudogenes in the mouse reference genome and associated strains. We compiled this by combining manual curation of over 10,000 pseudogenes with results from automatic annotation pipelines. Also, by comparing the human and mouse, we annotated 165 unitary pseudogenes in mouse, and 303 unitaries in human. We make all our annotation available through mouse.pseudogene.org. The overall mouse pseudogene repertoire (in the reference and strains) is similar to human in terms of overall size, biotype distribution (~80% processed/~20% duplicated) and top family composition (with many GAPDH and ribosomal pseudogenes). However, notable differences arise in the pseudogene age distribution, with multiple retro-transpositional bursts in mouse evolutionary history and only one in human. Furthermore, in each strain about a fifth of the pseudogenes are unique, reflecting strain-specific functions and evolution. Additionally, we find that ~15% of the pseudogenes are transcribed, a fraction similar to that for human, and that pseudogene transcription exhibits greater tissue and strain specificity compared to protein-coding genes. Finally, we show that highly transcribed parent genes tend to give rise to processed pseudogenes.

genomics

Nearly all new protein-coding predictions in the CHESS database are not protein-coding

In a 2018 paper posted to bioRxiv, Pertea et al. presented the CHESS database, a new catalog of human gene annotations that includes 1,178 new protein-coding predictions. These are based on evidence of transcription in human tissues and homology to earlier annotations in human and other mammals. Here, we reanalyze the evidence used by CHESS, and find that nearly all protein-coding predictions are false positives. We find that 86% overlap transposons marked by RepeatMasker that are known to frequently result in false positive protein-coding predictions. More than half are homologous to only nine Alu-derived primate sequences corresponding to an erroneous and previously withdrawn Pfam protein domain. The entire set shows poor evolutionary conservation and PhyloCSF protein-coding evolutionary signatures indistinguishable from noncoding RNAs, indicating lack of protein-coding constraint. Only four predictions are supported by mass spectrometry evidence, and even those matches are inconclusive. Overall, the new protein-coding predictions are unsupported by any credible experimental or evolutionary evidence of function, result primarily from homology to genes incorrectly classified as protein-coding, and are unlikely to encode functional proteins.

genomics

Epigenomic and functional dynamics of human bone marrow myeloid differentiation to mature blood neutrophils

Neutrophils are short-lived blood cells that play a critical role in host defense against infections. To better comprehend neutrophil functions and their regulation, we provide a complete epigenetic and functional overview of their differentiation stages from bone marrow-residing progenitors to mature circulating cells. Integration of epigenetic and transcriptome dynamics reveals an enforced regulation of differentiation, through cellular functions such as: release of proteases, respiratory burst, cell cycle regulation and apoptosis. We observe an early establishment of the cytotoxic capability, whilst the signaling components that activate antimicrobial mechanisms are transcribed at later stages, outside the bone marrow, thus preventing toxic effects in the bone marrow niche. Altogether, these data reveal how the developmental dynamics of the epigenetic landscape orchestrate the daily production of large number of neutrophils required for innate host defense and provide a comprehensive overview of the epigenomes of differentiating human neutrophils.\n\nKey pointsO_LIDynamic acetylation enforces human neutrophil progenitor differentiation.\nC_LI\n\nO_LINeutrophils cytotoxic capability is established early at the (pro)myelocyte stage.\nC_LI\n\nO_LICoordinated signaling component expression prevents unwanted toxic effects to the bone marrow niche.\nC_LI

immunology

Multiple laboratory mouse reference genomes define strain specific haplotypes and novel functional loci

The most commonly employed mammalian model organism is the laboratory mouse. A wide variety of genetically diverse inbred mouse strains, representing distinct physiological states, disease susceptibilities, and biological mechanisms have been developed over the last century. We report full length draft de novo genome assemblies for 16 of the most widely used inbred strains and reveal for the first time extensive strain-specific haplotype variation. We identify and characterise 2,567 regions on the current Genome Reference Consortium mouse reference genome exhibiting the greatest sequence diversity between strains. These regions are enriched for genes involved in defence and immunity, and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. Several immune related loci, some in previously identified QTLs for disease response have novel haplotypes not present in the reference that may explain the phenotype. We used these genomes to improve the mouse reference genome resulting in the completion of 10 new gene structures, and 62 new coding loci were added to the reference genome annotation. Notably this high quality collection of genomes revealed a previously unannotated gene (Efcab3-like) encoding 5,874 amino acids, one of the largest known in the rodent lineage. Interestingly, Efcab3-like-/- mice exhibit severe size anomalies in four regions of the brain suggesting a mechanism of Efcab3-like regulating brain development.

genomics

Genome variation and conserved regulation identify genomic regions responsible for strain specific phenotypes in rat

The genomes of laboratory rat strains are characterised by a mosaic haplotype structure caused by their unique breeding history. These mosaic haplotypes have been recently mapped by extensive sequencing of key strains. Comparison of genomic variation between two closely related rat strains with different phenotypes has been proposed as an effective strategy for the discovery of candidate strain-specific regions involved in phenotypic differences.\n\nWe developed a method to prioritise strain-specific haplotypes by integrating genomic variation and genomic regulatory data predicted to be involved in specific phenotypes. To identify genomic regions associated with metabolic syndrome, a disorder of energy utilization and storage affecting several organ systems, we compared two Lyon rat strains, LH/Mav which is susceptible to MetS, and LL/Mav, which is susceptible to obesity as an intermediate MetS phenotype, with a third strain (LN/Mav) that is resistant to both MetS and obesity. Applying a novel metric, we ranked the identified strain-specific haplotypes using evolutionary conservation of the occupancy three liver-specific transcription factors (HNF4A, CEBPA, and FOXA1) in five rodents including rat.\n\nConsideration of regulatory information effectively identified regions with liver-associated genes and rat orthologues of human GWAS variants related to obesity and metabolic traits. We attempted to find possible causative variants and compared them with the candidate genes proposed by previous studies. In strain-specific regions with conserved regulation, we found a significant enrichment for published evidence to obesity--one of the metabolic symptoms shown by the Lyon strains--amongst the genes assigned to promoters with strain-specific variation.\n\nOur results show that the use of functional regulatory conservation is a potentially effective approach to select strain-specific genomic regions associated with phenotypic differences among Lyon rats and could be extended to other systems.

genomics

Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes

Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 to 6 MYA, but that are absent in the Hominidae. In fact, Hominidae show between four-and seven-fold lower rates of nucleotide change and feature turnover in both neutral and functional sequences suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. For example, recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli. This process resulted in thousands of novel, species-specific CTCF binding sites. Our results demonstrate that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.

genomics

A Standardized Framework For Representation Of Ancestry Data In Genomics Studies

BackgroundThe accurate description of ancestry is essential to interpret and integrate human genomics data, and to ensure that advances in the field of genomics benefit individuals from all ancestral backgrounds. However, there are no established guidelines for the consistent, unambiguous and standardized description of ancestry. To fill this gap, we provide a framework, designed for the representation of ancestry in GWAS data, but with wider application to studies and resources involving human subjects.\n\nResultHere we describe our framework and its application to the representation of ancestry data in a widely-used publically available genomics resource, the NHGRI-EBI GWAS Catalog. We present the first analyses of GWAS data using our ancestry categories, demonstrating the validity of the framework to facilitate the tracking of ancestry in big data sets. We exhibit the broader relevance and integration potential of our method by its usage to describe the well-established HapMap and 1000 Genomes reference populations. Finally, to encourage adoption, we outline recommendations for authors to implement when describing samples.\n\nConclusionsWhile the known bias towards inclusion of European ancestry individuals in GWA studies persists, African and Hispanic or Latin American ancestry populations contribute a disproportionately high number of associations, suggesting that analyses including these groups may be more effective at identifying new associations. We believe the widespread adoption of our framework will increase standardization of ancestry data, thus enabling improved analysis, interpretation and integration of human genomics data and furthering our understanding of disease.

genetics

Complexity and conservation of regulatory landscapes underlie evolutionary resilience of mammalian gene expression

To gain insight into how mammalian gene expression is controlled by rapidly evolving regulatory elements, we jointly analysed promoter and enhancer activity with downstream transcription levels in liver samples from twenty species. Genes associated with complex regulatory landscapes generally exhibit high expression levels that remain evolutionarily stable. While the number of regulatory elements is the key driver of transcriptional output and resilience, regulatory conservation matters: elements active across mammals most effectively stabilise gene expression. In contrast, recently-evolved enhancers typically contribute weakly, consistent with their high evolutionary plasticity. These effects are observed across the entire mammalian clade and robust to potential confounders, such as gene expression level. Overall, our results illuminate how the evolutionary stability of gene expression is profoundly entwined with both the number and conservation of surrounding promoters and enhancers.\n\nHighlightsO_LIGene expression levels and stability are linked to the number of elements in the regulatory landscape.\nC_LIO_LIConserved regulatory elements associate with tightly controlled, highly expressed genes.\nC_LIO_LIRecently evolved enhancers weakly influence gene expression, but promoters are similarly active regardless of conservation.\nC_LIO_LIThe interplay between complexity of the regulatory landscape and conservation of individual promoters and enhancers shapes gene expression in mammals.\nC_LI

genomics

Genetic variation and gene expression across multiple tissues and developmental stages in a non-human primate

By analyzing multi-tissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalogue of expression quantitative trait loci (eQTLs) in a non-human primate model. This catalogue contains more genome-wide significant eQTLs, per sample, than comparable human resources, and reveals sex and age-related expression patterns. Findings include a master regulatory locus that likely plays a role in immune function, and a locus regulating hippocampal long non-coding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.

genetics

Chromosome assembly of large and complex genomes using multiple references

Despite the rapid development of sequencing technologies, assembly of mammalian-scale genomes into complete chromosomes remains one of the most challenging problems in bioinformatics. To help address this difficulty, we developed Ragout, a reference-assisted assembly tool that now works for large and complex genomes. Taking one or more target assemblies (generated from an NGS assembler) and one or multiple related reference genomes, Ragout infers the evolutionary relationships between the genomes and builds the final assemblies using a genome rearrangement approach. Using Ragout, we transformed NGS assemblies of 15 different Mus musculus and one Mus spretus genomes into sets of complete chromosomes, leaving less than 5% of sequence unlocalized per set. Various benchmarks, including PCR testing and realigning of long PacBio reads, suggest only a small number of structural errors in the final assemblies, comparable with direct assembly approaches. Additionally, we applied Ragout to Mus caroli and Mus pahari genomes, which exhibit karyotype-scale variations compared to other genomes from the Muridae family. Chromosome color maps confirmed most large-scale rearrangements that Ragout detected.

bioinformatics

Ensembl Core Software Resources: storage and programmatic access for DNA sequence and genome annotation

The Ensembl software resources are a stable infrastructure to store, access and manipulate genome assemblies and their functional annotations. The Ensembl \"Core\" database and Application Programming Interface (API) was our first major piece of software infrastructure and remains at the centre of all of our genome resources. Since its initial design more than fifteen years ago, the number of publicly available genomic, transcriptomic and proteomic datasets has grown enormously, accelerated by continuous advances in DNA sequencing technology. Initially intended to provide annotation for the reference human genome, we have extended our framework to support the genomes of all species as well as richer assembly models. Cross-referenced links to other informatics resources facilitate searching our database with a variety of popular identifiers such as UniProt and RefSeq. Our comprehensive and robust framework storing a large diversity of genome annotations in one location serves as a platform for other groups to generate and maintain their own tailored annotation. Our databases and APIs are publicly available and all of our source code is released with a permissive Apache v2.0 licence at http://github.com/Ensembl.

bioinformatics