bioRxiv ScienceSearch

Biology subjects

Gorjanc, G.

Publications and source records attributed to Gorjanc, G..

14 recordsLinked to original sources

The effects of training population design on genomic prediction accuracy in wheat

Genomic selection offers several routes for increasing genetic gain or efficiency of plant breeding programs. In various species of livestock there is empirical evidence of increased rates of genetic gain from the use of genomic selection to target different aspects of the breeders equation. Accurate predictions of genomic breeding value are central to this and the design of training sets is in turn central to achieving sufficient levels of accuracy. In summary, small numbers of close relatives and very large numbers of distant relatives are expected to enable accurate predictions.\n\nTo quantify the effect of some of the properties of training sets on the accuracy of genomic selection in crops we performed an extensive field-based winter wheat trial. In summary, this trial involved the construction of 44 F2:4 bi- and triparental populations, from which 2992 lines were grown on four field locations and yield was measured. For each line, genotype data were generated for 25,000 segregating single nucleotide polymorphism markers. The overall heritability of yield was estimated to 0.65, and estimates within individual families ranged between 0.10 and 0.85. Within cross genomic prediction accuracies of yield BLUEs were 0.125 - 0.127 using two different cross-validation approaches, and generally increased with training set size. Using related crosses in training and validation sets generally resulted in higher prediction accuracies than using unrelated crosses. The results of this study emphasize the importance of the training set design in relation to the genetic material to which the resulting prediction model is to be applied.

genetics

Analysis of a large data set reveals haplotypes carrying putatively recessive lethal alleles with pleiotropic effects on economically important traits in beef cattle

BackgroundDeleterious recessive alleles can result in reduced economic performance in livestock in multiple ways in homozygous individuals: from early embryonic death, death soon after birth, to being non-lethal but causing reduced viability. While death is an easy phenotype to score, reduced viability is not as easy to identify. However, it can sometimes be observed as reduced artificial insemination (AI) conception rates, longer calving intervals, or higher hazard for live born animals.\n\nMethodsIn this paper, we searched for haplotypes carrying putatively recessive lethal alleles in 132,725 genotyped Irish beef cattle from five breeds: Aberdeen Angus, Charolais, Hereford, Limousin, and Simmental. We phased the genotypes in sliding windows along the genome and used five tests to identify haplotypes with absence of or reduced homozygosity. We then corroborated the identified haplotypes with reproduction records, indicating early embryonic death, and postnatal survival records. Finally, we assessed haplotype pleiotropy by estimating substitution effects on national estimates of breeding values for 15 economically important traits in beef production.\n\nResultsWe found support for three haplotypes with carrying putatively recessive lethal alleles. The haplotypes were located on chromosome 14 in Aberdeen Angus, chromosome 19 in Charolais and chromosome 16 in Simmental. Their population frequencies is 15.2%, 14.4%, and 8.8%, respectively. All of the haplotypes showed pleiotropic effects on economically important traits for beef production. Their allele substitution effects are {euro}3.23, {euro}1.47, and {euro}2.30 for the terminal index and -{euro}3.15, -{euro}0.75, and {euro}1.12 for the replacement index, where one standard deviations are {euro}18.32, {euro}22.54, and {euro}22.33 for terminal index and {euro}29.52, {euro}35.62, and {euro}30.97 for the replacement index. We identified ZFAT as the candidate gene for lethality in Aberdeen Angus, several candidate genes for the Simmental haplotype, and no candidate genes for the Charolais haplotype.\n\nConclusionsWe analysed genotype, reproduction, survival, and production data to discover haplotypes carrying putatively recessive lethal alleles in Irish beef cattle. We found support for three haplotypes. All three haplotypes have pleiotropic effects on economically important traits in beef production.

genetics

Impact of index hopping and bias towards the reference allele on accuracy of genotype calls from low-coverage sequencing

BackgroundInherent sources of error and bias that affect the quality of the sequence data include index hopping and bias towards the reference allele. The impact of these artefacts is likely greater for low-coverage data than for high-coverage data because low-coverage data has scant information and standard tools for processing sequence data were designed for high-coverage data. With the proliferation of cost-effective low-coverage sequencing there is a need to understand the impact of these errors and bias on resulting genotype calls.\n\nResultsWe used a dataset of 26 pigs sequenced both at 2x with multiplexing and at 30x without multiplexing to show that index hopping and bias towards the reference allele due to alignment had little impact on genotype calls. However, pruning of alternative haplotypes supported by a number of reads below a predefined threshold, a default and desired step for removing potential sequencing errors in high-coverage data, introduced an unexpected bias towards the reference allele when applied to low-coverage data. This bias reduced best-guess genotype concordance of low-coverage sequence data by 19.0 absolute percentage points.\n\nConclusionsWe propose a simple pipeline to correct this bias and we recommend that users of low-coverage sequencing be wary of unexpected biases produced by tools designed for high-coverage sequencing.

bioinformatics

Sequence variability, constraint and selection in the CD163 gene in pigs

BackgroundIn this paper, we investigate sequence variability, evolutionary constraint, and selection on the CD163 gene in pigs. The pig CD163 gene is required for infection by porcine reproductive and respiratory syndrome virus (PRRSV), a serious pathogen with major impact on pig production.\n\nResultsWe used targeted pooled sequencing of the exons of CD163 to detect sequence variants in 35,000 pigs of diverse genetic backgrounds and search for potential knock-out variants. We then used whole genome sequence data from three pig lines to calculate a variant intolerance score, which measures the tolerance of genes to protein coding variation, a selection test on protein coding variation over evolutionary time, and haplotype diversity statistics to detect recent selective sweeps during breeding.\n\nConclusionsWe performed a deep survey of sequence variation in the CD163 gene in domestic pigs. We found no potential knock-out variants. CD163 was moderately intolerant to variation, and showed evidence of positive selection in the lineage leading up to the pig, but no evidence of selective sweeps during breeding.

genomics

Removal of alleles by genome editing -- RAGE against the deleterious load

BackgroundIn this paper, we simulate deleterious load in an animal breeding program, and compare the efficiency of genome editing and selection for decreasing load. Deleterious variants can be identified by bioinformatics screening methods that use sequence conservation and biological prior information about protein function. Once deleterious variants have been identified, how can they be used in breeding?\n\nResultsWe simulated a closed animal breeding population subject to both natural selection against deleterious load and artificial selection for a quantitative trait representing the breeding goal. Deleterious load was polygenic and due to either codominant or recessive variants. We compared strategies for removal of deleterious alleles by genome editing (RAGE) to selection against carriers. Each strategy varied in how animals and variants were prioritized for editing or selection.\n\nConclusionsGenome editing of deleterious alleles reduces deleterious load, but requires simultaneous editing of multiple deleterious variants in the same sire to be effective when deleterious variants are recessive. In the short term, selection against carriers is a possible alternative to genome editing when variants are recessive. The dominance of deleterious variants affects both the efficiency of genome editing and selection against carriers, and which variant prioritization strategy is the most efficient. Our results suggest that in the future, there is the potential to use RAGE against deleterious load in animal breeding.

genetics

A heuristic method for fast and accurate phasing and imputation of single nucleotide polymorphism data in bi-parental plant populations

This paper presents a new heuristic method for phasing and imputation of genomic data in diploid plant species. Our method, called AlphaPlantImpute, explicitly leverages features of plant breeding programs to maximise the accuracy of imputation. The features are a small number of parents, which can be inbred and usually have high-density genomic data, and few recombinations separating parents and focal individuals genotyped at low-density (i.e. descendants that are the imputation targets). AlphaPlantImpute works roughly in three steps. First, it identifies informative low-density genotype markers in parents. Second, it tracks the inheritance of parental alleles and haplotypes to focal individuals at informative markers. Finally, it uses this low-density information as anchor points to impute focal individuals to high-density.\n\nWe tested the imputation accuracy of AlphaPlantImpute in simulated bi-parental populations across different scenarios. We also compared its accuracy to existing software called PlantImpute. In general, AlphaPlantImpute had better or equal imputation accuracy as PlantImpute. The computational time and memory requirements of AlphaPlantImpute were tiny compared to PlantImpute. For example, accuracy of imputation was 0.96 for a scenario where both parents were inbred and genotyped at 25,000 markers per chromosome and a focal F2 individual was genotyped with 50 markers per chromosome. The maximum memory requirement for this scenario was 0.08 GB and took 37 seconds to complete.

genomics

Genomic prediction using individual-level data and summary statistics from multiple populations

This study presents a method for genomic prediction that uses individual-level data and summary statistics from multiple populations. Genome-wide markers are nowadays widely used to predict complex traits, and genomic prediction using multi-population data is an appealing approach to achieve higher prediction accuracies. However, sharing of individual-level data across populations is not always possible. We present a method that enables integration of summary statistics from separate analyses with the available individual-level data. The data can either consist of individuals with single or multiple (weighted) phenotype records per individual. We developed a method based on a hypothetical joint analysis model and absorption of population specific information. We show that population specific information is fully captured by estimated allele substitution effects and the accuracy of those estimates, i.e. the summary statistics. The method gives identical result as the joint analysis of all individual-level data when complete summary statistics are available. We provide a series of easy-to-use approximations that can be used when complete summary statistics are not available or impractical to share. Simulations show that approximations enables integration of different sources of information across a wide range of settings yielding accurate predictions. The method can be readily extended to multiple-traits. In summary, the developed method enables integration of genome-wide data in the individual-level or summary statistics form from multiple populations to obtain more accurate estimates of allele substitution effects and genomic predictions.

genomics

Parentage assignment with low density array data and low coverage sequence data

In this paper we evaluate using genotype-by-sequencing (GBS) data to perform parentage assignment in lieu of traditional array data. The use of GBS data raises two issues: First, for low-coverage GBS data, it may not be possible to call the genotype at many loci, a critical first step for detecting opposing homozygous markers. Second, the amount of sequencing coverage may vary across individuals, making it challenging to directly compare the likelihood scores between putative parents. To address these issues we extend the probabilistic framework of Huisman (2017) and evaluate putative parents by comparing their (potentially noisy) genotypes to a series of proposal distributions. These distributions describe the expected genotype probabilities for the relatives of an individual. We assign putative parents as a parent if they are classified as a parent (as opposed to e.g., an unrelated individual), and if the assignment score passes a threshold. We evaluated this method on simulated data and found that (1) high-coverage GBS data performs similarly to array data and requires only a small number of markers to correctly assign parents and (2) low-coverage GBS data (as low as 0.1x) can also be used, provided that it is obtained across a large number of markers. When analysing the low-coverage GBS data, we also found a high number of false positives if the true parent is not contained within the list of candidate parents, but that this false positive rate can be greatly reduced by hand tuning the assignment threshold. We provide this parentage assignment method as a standalone program called AlphaAssign.

genetics

AlphaMate: a program for optimising selection, maintenance of diversity, and mate allocation in breeding programs

SummaryAlphaMate is a flexible program that optimises selection, maintenance of genetic diversity, and mate allocation in breeding programs. It can be used in animal and cross- and self-pollinating plant populations. These populations can be subject to selective breeding or conservation management. The problem is formulated as a multi-objective optimisation of a valid mating plan that is solved with an evolutionary algorithm. A valid mating plan is defined by a combination of mating constraints (the number of matings, the maximal number of parents, the minimal/equal/maximal number of contributions per parent, or allowance for selfing) that are gender specific or generic. The optimisation can maximize genetic gain, minimize group coancestry, minimize inbreeding of individual matings, or maximize genetic gain for a given increase in group coancestry or inbreeding. Users provide a list of candidate individuals with associated gender and selection criteria information (if applicable) and coancestry matrix. Selection criteria and coancestry matrix can be based on pedigree or genome-wide markers. Additional individual or mating specific information can be included to enrich optimisation objectives. An example of rapid recurrent genomic selection in wheat demonstrates how AlphaMate can double the efficiency of converting genetic diversity into genetic gain compared to truncation selection. Another example demonstrates the use of genome editing to expand the gain-diversity frontier.\n\nAvailabilityExecutable versions of AlphaMate for Windows, Mac, and Linux platforms are available at http://www.alpha-genes.roslin.ed.ac.uk/AlphaMate\n\nContactgregor.gorjanc@roslin.ed.ack.uk

genetics

Optimal cross selection for long-term genetic gain in two-part programs with rapid recurrent genomic selection

This study evaluates optimal cross selection for balancing selection and maintenance of genetic diversity in two-part plant breeding programs with rapid recurrent genomic selection. The two-part program reorganizes a conventional breeding program into population improvement component with recurrent genomic selection to increase the mean of germplasm and product development component with standard methods to develop new lines. Rapid recurrent genomic selection has a large potential, but is challenging due to genotyping costs or genetic drift. Here we simulate a wheat breeding program for 20 years and compare optimal cross selection against truncation selection in the population improvement with one to six cycles per year. With truncation selection we crossed a small or a large number of parents. With optimal cross selection we jointly optimised selection, maintenance of genetic diversity, and cross allocation with AlphaMate program. The results show that the two-part program with optimal cross selection delivered the largest genetic gain that increased with the increasing number of cycles. With four cycles per year optimal cross selection had 78% (15%) higher long-term genetic gain than truncation selection with a small (large) number of parents. Higher genetic gain was achieved through higher efficiency of converting genetic diversity into genetic gain; optimal cross selection quadrupled (doubled) efficiency of truncation selection with a small (large) number of parents. Optimal cross selection also reduced the drop of genomic selection accuracy due to the drift between training and prediction populations. In conclusion, optimal cross-selection enables optimal management and exploitation of population improvement germplasm in two-part programs.\n\nKey messageOptimal cross selection increases long-term genetic gain of two-part programs with rapid recurrent genomic selection. It achieves this by optimising efficiency of converting genetic diversity into genetic gain through reducing the loss of genetic diversity and reducing the drop of genomic prediction accuracy with rapid cycling.

genetics

Hybrid peeling for fast and accurate calling, phasing, and imputation with sequence data of any coverage in pedigrees

In this paper we extend multi-locus iterative peeling to be a computationally efficient method for calling, phasing, and imputing sequence data of any coverage in small or large pedigrees. Our method, called hybrid peeling, uses multi-locus iterative peeling to estimate shared chromosome segments between parents and their offspring, and then uses single-locus iterative peeling to aggregate genomic information across multiple generations. Using a synthetic dataset, we first analysed the performance of hybrid peeling for calling and phasing alleles in disconnected families, families which contained only a focal individual and its parents and grandparents. Second, we analysed the performance of hybrid peeling for calling and phasing alleles in the context of the full pedigree. Third, we analysed the performance of hybrid peeling for imputing whole genome sequence data to the remaining individuals in the population. We found that hybrid peeling substantially increase the number of genotypes that were called and phased by leveraging sequence information on related individuals. The calling rate and accuracy increased when the full pedigree was used compared to a reduced pedigree of just parents and grandparents. Finally, hybrid peeling accurately imputed whole genome sequence information to non-sequenced individuals. We believe that this algorithm will enable the generation of low cost and high accuracy whole genome sequence data in many pedigreed populations. We are making this algorithm available as a standalone program called AlphaPeel.

genetics

Assessment of the performance of different hidden Markov models for imputation in animal breeding

In this paper we review the performance of various hidden Markov model-based imputation methods in animal breeding populations. Traditionally, heuristic-based imputation methods have been used for imputation in large animal populations due to their computational efficiency, scalability, and accuracy. However, recent advances in the area of human genetics have increased the ability of probabilistic hidden Markov model methods to perform accurate phasing and imputation in large populations. These advances may enable these methods to be useful for routine use in large animal populations. To test this, we evaluate here the accuracy and computational cost of several methods in a series of simulated populations and a real animal population. We first tested single-step (diploid) imputation, which performs both phasing and imputation. Then we tested pre-phasing followed by haploid imputation. We tested four diploid imputation methods (fastPHASE, Beagle v4.0, IMPUTE2, and MaCH), three phasing methods, (SHAPEIT2, HAPI-UR, and Eagle2), and three haploid imputation methods (IMPUTE2, Beagle v4.1, and minimac3). We found that performing pre-phasing and haploid imputation was faster and more accurate than diploid imputation. In particular, we found that pre-phasing with Eagle2 or HAPI-UR and imputing with minimac3 or IMPUTE2 gave the highest accuracies in both simulated and real data.

genetics

A strategy to exploit surrogate sire technology in livestock breeding programs

In this work, we performed simulations to develop and test a strategy for exploiting surrogate sire technology in animal breeding programs. Surrogate sire technology allows the creation of males that lack their own germline cells, but have transplanted spermatogonial stem cells from donor males. With this technology, a single elite male donor could give rise to huge numbers of progeny, potentially as much as all the production animals in a particular time period.\n\nOne hundred replicates of various scenarios were performed. Scenarios followed a common overall structure but differed in the strategy used to identify elite donors and how these donors were used in the product development part.\n\nThe results of this study showed that using surrogate sire technology would significantly increase the genetic merit of commercial sires, by as much as 6.5 to 9.2 years worth of genetic gain compared to a conventional breeding program. The simulations suggested that a strategy involving three stages (an initial genomic test followed by two subsequent progeny tests) was the most effective of all the strategies tested.\n\nThe use of one or a handful of elite donors to generate the production animals would be very different to current practice. While the results demonstrate the great potential of surrogate sire technology there are considerable risks but also other opportunities. Practical implementation of surrogate sire technology would need to account for these.

genetics

A method for allocating low-coverage sequencing resources by targeting haplotypes rather than individuals

BackgroundThis paper describes a heuristic method for allocating low-coverage sequencing resources by targeting haplotypes rather than individuals. Low-coverage sequencing assembles high-coverage sequence information for every individual by accumulating data from the genome segments that they share with many other individuals into consensus haplotypes. Deriving the consensus haplotypes accurately is critical for achieving a high phasing and imputation accuracy. In order to enable accurate phasing and imputation of sequence information for the whole population we allocate the available sequencing resources among individuals with existing phased genomic data by targeting the sequencing coverage of their haplotypes.\n\nResultsOur method, called AlphaSeqOpt, prioritizes haplotypes using a score function that is based on the frequency of the haplotypes in the sequencing set relative to the target coverage. AlphaSeqOpt has two steps: (1) selection of an initial set of individuals by iteratively choosing the individuals that have the maximum score conditional to the current set, and (2) refinement of the set through several rounds of exchanges of individuals. AlphaSeqOpt is very effective for distributing a fixed amount of sequencing resources evenly across haplotypes, which results in a reduction of the proportion of haplotypes that are sequenced below the target coverage. AlphaSeqOpt can provide a greater proportion of haplotypes sequenced at the target coverage by sequencing less individuals, as compared with other methods that use a score function based on the haplotypes population frequency. A refinement of the initially selected set can provide a larger more diverse set with more unique individuals, which is beneficial in the context of low-coverage sequencing. We extend the method with an approach to filter rare haplotypes based on their flanking haplotypes, so that only those that are likely to derive from a recombination event are targeted.\n\nConclusionsWe present a method for allocating sequencing resources so that a greater proportion of haplotypes are sequenced at a coverage that is sufficiently high for population-based imputation with low-coverage sequencing. The haplotype score function, the refinement step, and the new approach of filtering rare haplotypes make AlphaSeqOpt more effective for that purpose than methods reported previously for reducing sequencing redundancy.

genomics