bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Improving the Efficiency of Genomic Selection in Chinese Simmental beef cattle

Genomic selection is an accurate and efficient method of estimating genetic merits by using high-density genome-wide single nucleotide polymorphisms (SNPs).In this study, we investigate an approach to increase the efficiency of genomic prediction by using genome-wide markers. The approach is a feature selection based on genomic best linear unbiased prediction (GBLUP),which is a statistical method used to predict breeding values using SNPs for selection in animal and plant breeding. The objective of this study is the choice of kinship matrix for genomic best linear unbiased prediction (GBLUP).The G-matrix is using the information of genome-wide dense markers. We compare three kinds of kinships based on different combinations of centring and scaling of marker genotypes. And find a suitable kinship approach that adjusts for the resource population of Chinese Simmental beef cattle. Single nucleotide polymorphism (SNPs) can be used to estimate kinship matrix and individual inbreeding coefficients more accurately. So in our research a genomic relationship matrix was developed for 1059 Chinese Simmental beef cattle using 640000 single nucleotide polymorphisms and breeding values were estimated using phenotypes about Carcass weight and Sirloin weight. The number of SNPs needed to accurately estimate a genomic relationship matrix was evaluated in this population. Another aim of this study was to optimize the selection of markers and determine the required number of SNPs for estimation of kinship in the Chinese Simmental beef cattle.\n\nWe find that the feature selection of GBLUP using Xus and the Astle and Baldings kinships model performed similarly well, and were the best-performing methods in our study. Inbreeding and kinship matrix can be estimated with high accuracy using [≥]12,000s in Chinese Simmental beef cattle.

Genomics↗

Utilization of high throughput genome sequencing technology for large scale single nucleotide polymorphism discovery in red deer and Canadian elk

Deer farming is a significant international industry. For genetic improvement, using genomic tools, an ordered array of DNA variants and associated flanking sequence across the genome is required. This work reports a comparative assembly of the deer genome and subsequent DNA variant identification. Next generation sequencing combined with an existing bovine reference genome enabled the deer genome to be assembled sufficiently for large-scale SNP discovery. In total, 28 Gbp of sequence data were generated from seven Cervus elaphus (European red deer and Canadian elk) individuals. After aligning sequence to the bovine reference genome build UMD 3.0 and binning reads into one Mbp groups; reads were assembled and analyzed for SNPs. Greater than 99% of the non-repetitive fraction of the bovine genome was covered by deer chromosomal scaffolds. We identified 1.8 million SNPs meeting Illumina InfiniumII SNP chip technical threshold. Markers on the published Red x Pere David deer linkage map were aligned to both UMD3.0 and the new deer chromosomal scaffolds. This enabled deer linkage groups to be assigned to deer chromosomal scaffolds, although the mapping locations remain based on bovine order. Genotyping of 270 SNPs on a Sequenom MS system showed that 88% of SNPs identified could be amplified. Also, inheritance patterns showed no evidence of departure from Hardy-Weinberg equilibrium. A comparative assembly of the deer genome, alignment with existing deer genetic linkage groups and SNP discovery has been successfully completed and validated facilitating application of genomic technologies for subsequent deer genetic improvement.

Genomics↗

Genome Sequences of Populus tremula Chloroplast and Mitochondrion: Implications for Holistic Poplar Breeding

Complete Populus genome sequences are available for the nucleus (P. trichocarpa; section Tacamahaca) and for chloroplasts (seven species), but not for mitochondria. Here, we provide the complete genome sequences of the chloroplast and the mitochondrion for the clones P. tremula W52 and P. tremula x P. alba 717-1B4 (section Populus). The organization of the chloroplast genomes of both Populus clones is described. A phylogenetic tree constructed from all available complete chloroplast DNA sequences of Populus was not congruent with the assignment of the related species to different Populus sections. In total, 3,024 variable nucleotide positions were identified among all compared Populus chloroplast DNA sequences. The 5-prime part of the LSC from trnH to atpA showed the highest frequency of variations. The variable positions included 163 positions with SNPs allowing for differentiating the two clones with P. tremula chloroplast genomes (W52 717-1B4) from the other seven Populus individuals. These potential P. tremula-specific SNPs were displayed as a whole-plastome barcode on the P. tremula W52 chloroplast DNA sequence. Three of these SNPs and one InDel in the trnH-psbA linker were successfully validated by Sanger sequencing in an extended set of Populus individuals. The complete mitochondrial genome sequence of P. tremula is the first in the family of Salicaceae. The mitochondrial genomes of the two clones are 783,442 bp (W52) and 783,513 bp (717-1B4) in size, structurally very similar and organized as single circles. DNA sequence regions with high similarity to the W52 chloroplast sequence account for about 2% of the W52 mitochondrial genome. The mean SNP frequency was found to be nearly six fold higher in the chloroplast than in the mitochondrial genome when comparing 717-1B4 with W52. The availability of the genomic information of all three DNA-containing cell organelles will allow a holistic approach in poplar molecular breeding in the future.

Genomics↗

Computational Pan-Genomics: Status, Promises and Challenges

Many disciplines, from human genetics and oncology to plant breeding, microbiology and virology, commonly face the challenge of analyzing rapidly increasing numbers of genomes. In case of Homo sapiens, the number of sequenced genomes will approach hundreds of thousands in the next few years. Simply scaling up established bioinformatics pipelines will not be sufficient for leveraging the full potential of such rich genomic datasets. Instead, novel, qualitatively different computational methods and paradigms are needed. We will witness the rapid extension of computational pan-genomics, a new sub-area of research in computational biology. In this paper, we generalize existing definitions and understand a pan-genome as any collection of genomic sequences to be analyzed jointly or to be used as a reference. We examine already available approaches to construct and use pan-genomes, discuss the potential benefits of future technologies and methodologies, and review open challenges from the vantage point of the above-mentioned biological disciplines. As a prominent example for a computational paradigm shift, we particularly highlight the transition from the representation of reference genomes as strings to representations as graphs. We outline how this and other challenges from different application domains translate into common computational problems, point out relevant bioinformatics techniques and identify open problems in computer science. With this review, we aim to increase awareness that a joint approach to computational pan-genomics can help address many of the problems currently faced in various domains.

Genomics↗

Snapshots of a shrinking partner: Genome reduction in Serratia symbiotica

Genome reduction is pervasive among maternally-inherited endosymbiotic organisms, from bacteriocyte- to gut-associated ones. This genome erosion is a step-wise process in which once free-living organisms evolve to become obligate associates, thereby losing non-essential or redundant genes/functions. Serratia symbiotica (Gammaproteobacteria), a secondary endosymbiont present in many aphids (Hemiptera: Aphididae), displays various characteristics that make it a good model organism for studying genome reduction. While some strains are of facultative nature, others have established co-obligate associations with their respective aphid host and its primary endosymbiont (Buchnera). Furthermore, the different strains hold genomes of contrasting sizes and features, and have strikingly disparate cell shapes, sizes, and tissue tropism. Finally, genomes from closely related free-living Serratia marcescens are also available. In this study, we describe in detail the genome reduction process (from free-living to reduced obligate endosymbiont) undergone by S. symbiotica, and relate it to the stages of integration to the symbiotic system the different strains find themselves in. We establish that the genome reduction patterns observed in S. symbiotica follow those from other dwindling genomes, thus proving to be a good model for the study of the genome reduction process within a single bacterial taxon evolving in a similar biological niche (aphid-Buchnera).

Genomics↗

An improved genome assembly uncovers a prolific tandem repeat structure in Atlantic cod

Background: The first Atlantic cod (Gadus morhua) genome assembly published in 2011 was one of the early genome assemblies exclusively based on high-throughput 454 pyrosequencing. Since then, rapid advances in sequencing technologies have led to a multitude of assemblies generated for complex genomes, although many of these are of a fragmented nature with a significant fraction of bases in gaps. The development of long-read sequencing and improved software now enable the generation of more contiguous genome assemblies.\n\nResults: By combining data from Illumina, 454 and the longer PacBio sequencing technologies, as well as integrating the results of multiple assembly programs, we have created a substantially improved version of the Atlantic cod genome assembly. The sequence contiguity of this assembly is increased fifty-fold and the proportion of gap-bases has been reduced fifteen-fold. Compared to other vertebrates, the assembly contains an unusual high density of tandem repeats (TRs). Indeed, retrospective analyses reveal that gaps in the first genome assembly were largely associated with these TRs. We show that 21 % of the TRs across the assembly, 19 % in the promoter regions and 12 % in the coding sequences are heterozygous in the sequenced individual.\n\nConclusions: The inclusion of PacBio reads combined with the use of multiple assembly programs drastically improved the Atlantic cod genome assembly by successfully resolving long TRs. The high frequency of heterozygous TRs within or in the vicinity of genes in the genome indicate a considerable standing genomic variation in Atlantic cod populations, which is likely of evolutionary importance.

Genomics↗

The human functional genome defined by genetic diversity

Large scale efforts to sequence whole human genomes provide extensive data on the non-coding portion of the genome. We used variation information from 11,257 human genomes to describe the spectrum of sequence conservation in the population. We established the genome-wide variability for each nucleotide in the context of the surrounding sequence in order to identify departure from expectation at the population level (context-dependent conservation). We characterized the population diversity for functional elements in the genome and identified the coordination of conserved sequences of distal and cis enhancers, chromatin marks, promoters, coding and intronic regions. The most context-dependent conserved regions of the genome are associated with unique functional annotations and a genomic organization that spreads up to one megabase. Importantly, these regions are enriched by over 100-fold of non-coding pathogenic variants. This analysis of human genetic diversity thus provides a detailed view of sequence conservation, functional constraint and genomic organization of the human genome. Specifically, it identifies highly conserved non-coding sequences that are not captured by analysis of interspecies conservation and are greatly enriched in disease variants.

genomics↗

Pairwise comparisons are problematic when analyzing functional genomic data across species

AbstractThere is considerable interest in comparing functional genomic data across species. One goal of such work is to provide an integrated understanding of genome and phenotype evolution. Most comparative functional genomic studies have relied on multiple pairwise comparisons between species, an approach that does not incorporate information about the evolutionary relationships among species. The statistical problems that arise from not considering these relationships can lead pairwise approaches to the wrong conclusions, and are a missed opportunity to learn about biology that can only be understood in an explicit phylogenetic context. Here we examine two recently published studies that compare gene expression across species with pairwise methods, and find reason to question the original conclusions of both. One study interpreted pairwise comparisons of gene expression as support for the ortholog conjecture, the hypothesis that orthologs tend to be more similar than paralogs. The other study interpreted pairwise comparisons of embryonic gene expression across distantly related animals as evidence for a distinct evolutionary process that gave rise to phyla. In each study, distinct patterns of pairwise similarity among species were originally interpreted as evidence of particular evolutionary processes, but instead we find they reflect species relationships. These reanalyses concretely demonstrate the inadequacy of pairwise comparisons for analyzing functional genomic data across species. It will be critical to adopt phylogenetic comparative methods in future functional genomic work. Fortunately, phylogenetic comparative biology is also a rapidly advancing field with many methods that can be directly applied to functional genomic data.\n\nSignificanceComparisons of genome function between species are providing important insight into the evolutionary origins of diversity. Here we demonstrate that comparative functional genomics studies can come to the wrong conclusions if they do not take the relationships of species into account and instead rely on pairwise comparisons between species, as is common practice. We re-examined two previously published studies and found problems with pairwise comparisons that draw both their original conclusions into question. One study had found support for the ortholog conjecture and the other had concluded that the evolution of gene expression was different between animal phyla than within them. Our results demonstrate that to answer evolutionary questions about genome function, it is critical to consider evolutionary relationships.

genomics↗

The Genomic Health Of Ancient Hominins

The genomes of ancient humans, Neandertals, and Denisovans contain many alleles that influence disease risks. Using genotypes at 3180 disease-associated loci, we estimated the disease burden of 147 ancient genomes. After correcting for missing data, genetic risk scores were generated for nine disease categories and the set of all combined diseases. These genetic risk scores were used to examine the effects of different types of subsistence, geography, and sample age on the number of risk alleles in each ancient genome. On a broad scale, hereditary disease risks are similar for ancient hominins and modern-day humans, and the GRS percentiles of ancient individuals span the full range of what is observed in present day individuals. In addition, there is evidence that ancient pastoralists may have had healthier genomes than hunter-gatherers and agriculturalists. We also observed a temporal trend whereby genomes from the recent past are more likely to be healthier than genomes from the deep past. This calls into question the idea that modern lifestyles have caused genetic load to increase over time. Focusing on individual genomes, we find that the overall genomic health of the Altai Neandertal is worse than 97% of present day humans and that Otzi the Tyrolean Iceman had a genetic predisposition to gastrointestinal and cardiovascular diseases. As demonstrated by this work, ancient genomes afford us new opportunities to diagnose past human health, which has previously been limited by the quality and completeness of remains.

genomics↗

High-quality de novo genome assembly of the Dekkera bruxellensis UMY321 yeast isolate using Nanopore MinION sequencing

Genetic variation in natural populations represents the raw material for phenotypic diversity. Species-wide characterization of genetic variants is crucial to have a deeper insight into the genotype-phenotype relationship. With the advent of new sequencing strategies and more recently the release of long-read sequencing platforms, it is now possible to explore the genetic diversity of any non-model organisms, representing a fundamental resource for biological research. In the frame of population genomic surveys, a first step is evidently to obtain the complete sequence and high quality assembly of a reference genome. Here, we completely sequenced and assembled a reference genome of the non-conventional Dekkera bruxellensis yeast. While this species is a major cause of wine spoilage, it paradoxically contributes to the specific flavor profile of some Belgium beers. In addition, an extreme karyotype variability is observed across natural isolates, highlighting that D. bruxellensis genome is very dynamic. The whole genome of the D. bruxellensis UMY321 isolate was sequenced using a combination of Nanopore long-read and Illumina short-read sequencing data. We generated the most complete and contiguous de novo assembly of D. bruxellensis to date and obtained a first glimpse into the genomic variability within this species by comparing the sequences of several isolates. This genome sequence is therefore of high value for population genomic surveys and represents a reference to study genome dynamic in this yeast species.

genomics↗

Fantastic beasts and how to sequence them: genomic approaches for obscure model organisms.

Application of genomic approaches to \"obscure model organisms\" (OMOs), meaning species with little or no genomic resources, enables increasingly sophisticated studies of genomic basis of evolution, acclimatization and adaptation in real ecological contexts. Here, I highlight sequencing solutions and data handling techniques most suited for genomic analysis of OMOs.\n\nGlossary- Allele Frequency Spectrum, AFS (same as Site Frequency Spectrum, SFS): histogram of the number of segregating variants depending on their frequency in one or more populations.\n- Restriction site-Associated DNA (RAD) sequencing: family of diverse genotyping methods that sequence short fragments of the genome adjacent to recognition site(s) for specific restriction endonuclease(s).\n- Linkage Disequilibrium (LD): in this review, correlation of genotypes at a pair of markers across individuals.\n- LD block: typical distance between markers in the genome across which their genotypes remain correlated.\n- Genome scan: profiling of genotypes along the genome looking for unusual patterns. Often used to look for signatures of natural selection or introgression.\n- \"Denser-than-LD\" genotyping: genotyping of several polymorphic markers per LD block.\n- Highly contiguous reference: genome or transcriptome reference sequence containing the least amount of fragmentation.\n- Phased data: data showing which SNP alleles belong to the same homologous chromosome copy.\n- Cross-tissue gene expression analysis: looking for individual-specific shifts in gene expression detectable across multiple tissues. Such shifts are predominantly genetic in nature.

genomics↗

Using DNase Hi-C techniques to map global and local three-dimensional genome architecture at high resolution

The folding and three-dimensional (3D) organization of chromatin in the nucleus critically impacts genome function. The past decade has witnessed rapid advances in genomic tools for delineating 3D genome architecture. Among them, chromosome conformation capture (3C)-based methods such as Hi-C are the most widely used techniques for mapping chromatin interactions. However, traditional Hi-C protocols rely on restriction enzymes (REs) to fragment chromatin and are therefore limited in resolution. We recently developed DNase Hi-C for mapping 3D genome organization, which uses DNase I for chromatin fragmentation. DNase Hi-C overcomes RE-related limitations associated with traditional Hi-C methods, leading to improved methodological resolution. Furthermore, combining this method with DNA capture technology provides a high-throughput approach (targeted DNase Hi-C) that allows for mapping fine-scale chromatin architecture at exceptionally high resolution. Hence, targeted DNase Hi-C will be valuable for delineating the physical landscapes of cis-regulatory networks that control gene expression and for characterizing phenotype-associated chromatin 3D signatures. Here, we provide a detailed description of method design and step-by-step working protocols for these two methods.\n\nHighlightsO_LIDNase Hi-C, a method for comprehensive mapping of chromatin contacts on a whole-genome scale, is based on random chromatin fragmentation by DNase I digestion instead of sequence-specific restriction enzyme (RE) digestion.\nC_LIO_LITargeted DNase Hi-C, which combines DNase Hi-C with DNA capture technology, is a high-throughput method for mapping fine-scale chromatin architecture of genomic loci of interest at a resolution comparable to that of genomic annotations of functional elements.\nC_LIO_LIDNase Hi-C and targeted DNase Hi-C provide the first high-throughput way to overcome the RE-digestion-associated resolution limit of 3C-based methods.\nC_LIO_LIStep-by-step whole-genome and targeted DNase Hi-C protocols for mapping global and local 3D genome architecture, respectively, are described.\nC_LI

genomics↗

UniProt Genomic Mapping for Deciphering Functional Effects of Missense Variants

Understanding the association of genetic variation with its functional consequences in proteins is essential for the interpretation of genomic data and identifying causal variants in diseases. Integration of protein function knowledge with genome annotation can assist in rapidly comprehending genetic variation within complex biological processes. Here, we describe mapping UniProtKB human sequences and positional annotations such as active sites, binding sites, and variants to the human genome (GRCh38) and the release of a public genome track hub for genome browsers. To demonstrate the power of combining protein annotations with genome annotations for functional interpretation of variants, we present specific biological examples in disease-related genes and proteins. Computational comparisons of UniProtKB annotations and protein variants with ClinVar clinically annotated SNP data show that 32% of UniProtKB variants co-locate with 8% of ClinVar SNPs. The majority of co-located UniProtKB disease-associated variants (86%) map to pathogenic ClinVar SNPs. UniProt and ClinVar are collaborating to provide a unified clinical variant annotation for genomic, protein and clinical researchers. The genome track hubs, and related UniProtKB files, are downloadable from the UniProt FTP site and discoverable as public track hubs at the UCSC and Ensembl genome browsers.

genomics↗

A critical comparison of technologies for a plant genome sequencing project

A high quality genome sequence of your model organism is an essential starting point for many studies. Old clone based methods are slow and expensive, whereas faster, cheaper short read only assemblies can be incomplete and highly fragmented, which minimises their usefulness. The last few years have seen the introduction of many new technologies for genome assembly. These new technologies and new algorithms are typically benchmarked on microbial genomes or, if they scale appropriately, human. However, plant genomes can be much more repetitive and larger than human, and plant biology makes obtaining high quality DNA free from contaminants difficult. Reflecting their challenging nature we observe that plant genome assembly statistics are typically poorer than for vertebrates. Here we compare Illumina short read, PacBio long read, 10x Genomics linked reads, Dovetail Hi-C and BioNano Genomics optical maps, singly and combined, in producing high quality long range genome assemblies of the potato species S. verrucosum. We benchmark the assemblies for completeness and accuracy, as well as DNA, compute requirements and sequencing costs. We expect our results will be helpful to other genome projects, and that these datasets will be used in benchmarking by assembly algorithm developers.

genomics↗

A highly contiguous genome for the Golden-fronted Woodpecker (Melanerpes aurifrons) via a hybrid Oxford Nanopore and short read assembly

BackgroundWoodpeckers are found in nearly every part of the world, absent only from Antarctica, Australasia, and Madagascar. Woodpeckers have been important for studies of biogeography, phylogeography, and macroecology. Woodpeckers hybrid zones are often studied to understand the dynamics of introgression between bird species. Notably, woodpeckers are gaining attention for their enriched levels of transposable elements (TEs) relative to most other birds. This enrichment of TEs may have substantial effects on woodpecker molecular evolution. The Golden-fronted Woodpecker (Melanerpes aurifrons) is a member of the largest radiation of New World woodpeckers. However, comparative studies of woodpecker genomes are hindered by the fact that no high-contiguity genome exists for any woodpecker species. FindingsUsing hybrid assembly methods that combine long-read Oxford Nanopore and short-read Illumina sequencing data, we generated a highly contiguous genome assembly for the Golden-fronted Woodpecker. The final assembly is 1.31 Gb and comprises 441 contigs plus a full mitochondrial genome. Half of the assembly is represented by 28 contigs (contig N50), each of these contigs is at least 16 Mb in size (contig L50). High recovery (92.6%) of bird-specific BUSCO genes suggests our assembly is both relatively complete and relatively accurate. Accuracy is also demonstrated by the recovery of a putatively error-free mitochondrial genome. Over a quarter (25.8%) of the genome consists of repetitive elements, with 287 Mb (21.9%) of those elements assignable to the CR1 superfamily of transposable elements, the highest proportion of CR1 repeats reported for any bird genome to date. ConclusionOur assembly provides a useful tool for comparative studies of molecular evolution and genomics in woodpeckers and allies, a group emerging as important for studies on the role that TEs may play in avian evolution. Additionally, the sequencing and bioinformatic resources used to generate this assembly were relatively low-cost and should provide a direction for the development of high-quality genomes for future studies of animal biodiversity.

genomics↗

Jekyll or Hyde? The genome (and more) of Nesidiocoris tenuis, a zoophytophagous predatory bug that is both a biological control agent and a pest

Nesidiocoris tenuis (Reuter) is an efficient predatory biological control agent used throughout the Mediterranean Basin in tomato crops but regarded as a pest in northern European countries. Belonging to the family Miridae, it is an economically important insect yet very little is known in terms of genetic information - no published genome, population studies, or RNA transcripts. It is a relatively small and long-lived diploid insect, characteristics that complicate genome sequencing. Here, we circumvent these issues by using a linked-read sequencing strategy on a single female N. tenuis. From this, we assembled the 355 Mbp genome and delivered an ab initio, homology-based, and evidence-based annotation. Along the way, the bacterial "contamination" was removed from the assembly, which also revealed potential symbionts. Additionally, bacterial lateral gene transfer (LGT) candidates were detected in the N. tenuis genome. The complete gene set is composed of 24,688 genes; the associated proteins were compared to other hemipterans (Cimex lectularis, Halyomorpha halys, and Acyrthosiphon pisum), resulting in an initial assessment of unique and shared protein clusters. We visualised the genome using various cytogenetic techniques, such as karyotyping, CGH and GISH, indicating a karyotype of 2n=32 with a male-heterogametic XX/XY system. Additional analyses include the localization of 18S rDNA and unique satellite probes via FISH techniques. Finally, population genomics via pooled sequencing further showed the utility of this genome. This is one of the first mirid genomes to be released and the first of a mirid biological control agent, representing a step forward in integrating genome sequencing strategies with biological control research.

genomics↗

Chromosome-Scale Genome Assembly Provides Insights into Speciation of Allotetraploid and Massive Biomass Accumulation of Elephant Grass (Pennisetum purpureum Schum.)

Elephant grass (Pennisetum purpureum Schum., AABB, 2n=4x=28), which is characterized as robust growth and high biomass, and widely distributed in tropical and subtropical areas globally, is an important forage, biofuels and industrial plant. We sequenced its allopolyploid genome and assembled 2.07 Gb (96.88%) into A and B sub-genomes of 14 chromosomes with scaffold N50 of 8.47 Mb. A total of 38,453 and 36,981 genes were annotated in A and B sub-genomes, respectively. A phylogenetic analysis with species in Pennisetum identified that the speciation of the allotetraploid occurred approximately 15 MYA after the divergence between S.italica and P. glaucum. Double whole-genome duplication (WGD) and polyploidization events resulted in large scale gene expansion, especially in the key steps of growth and biomass accumulation. Integrated transcriptome profiling revealed the functional differentiation between sub-genomes; A sub-genome contributed more to plant growth, development and photosynthesis whereas B sub-genome primarily offered functions of effective transportation and resistance to stimulation. The results uncovered enhanced cellulose and lignin biosynthesis pathways with 645 and 666 genes expanded in A and B sub-genomes, respectively. Our findings provided deep insights into the speciation and genetic basis of fast growth and high biomass accumulation in the species. The genetic, genomic, and transcriptomic resources generated in this study will pave the way for further domestication and selection of these economical species and making them more adaptive to industrial utilization.

genomics↗

Evolutionarily conserved non-protein-coding regions in the chicken genome harbor functionally important variation

The availability of genomes for many species has advanced our understanding of the non-protein-coding fraction of the genome. Comparative genomics has proven to be an invaluable approach for the systematic, genome-wide identification of conserved non-protein-coding elements (CNEs). However, for many non-mammalian model species, including chicken, our capability to interpret the functional importance of variants overlapping CNEs has been limited by current genomic annotations, which rely on a single information type (e.g. conservation). We here studied CNEs in chicken using a combination of population genomics and comparative genomics. To investigate the functional importance of variants found in CNEs we develop a ch(icken) Combined Annotation-Dependent Depletion (chCADD), a variant effect prediction tool first introduced for humans and later on for mouse and pig. We show that 73 Mb of the chicken genome has been conserved across more than 280 million years of vertebrate evolution. The vast majority of the conserved elements are in non-protein-coding regions, which display SNP densities and allele frequency distributions characteristic of genomic regions constrained by purifying selection. By annotating SNPs with the chCADD score we are able to pinpoint specific subregions of the CNEs to be of higher functional importance, as supported by SNPs found in these subregions are associated with known disease genes in humans, mice, and rats. Taken together, our findings indicate that CNEs harbor variants of functional significance that should be object of further investigation along with protein-coding mutations. We therefore anticipate chCADD to be of great use to the scientific community and breeding companies in future functional studies in chicken.

genomics↗