bioRxiv Science⌕ Search

Biology subjects

Long, E. M.

Publications and source records attributed to Long, E. M..

2 recordsLinked to original sources

Utilizing Evolutionary Conservation to Detect Deleterious Mutations and Improve Genomic Prediction in Cassava

Cassava (Manihot esculenta) is an annual root crop which provides the major source of calories for over half a billion people around the world. Since its domestication [~]10,000 years ago, cassava has been largely clonally propagated through stem cuttings. Minimal sexual recombination has led to an accumulation of deleterious mutations made evident by heavy inbreeding depression. To locate and characterize these deleterious mutations, and to measure selection pressure across the cassava genome, we aligned 52 related Euphorbiaceae and other related species representing millions of years of evolution. With single base-pair resolution of genetic conservation, we used protein structure models, amino acid impact, and evolutionary conservation across the Euphorbiaceae to estimate evolutionary constraint. With known deleterious mutations, we aimed to improve genomic evaluations of plant performance through genomic prediction. We first tested this hypothesis through simulation utilizing multi-kernel GBLUP to predict simulated phenotypes across separate populations of cassava. Simulations showed a sizable increase of prediction accuracy when incorporating functional variants in the model when the trait was determined by <100 quantitative trait loci (QTL). Utilizing deleterious mutations and functional weights informed through evolutionary conservation, we saw improvements in genomic prediction accuracy that were dependent on trait and prediction. We showed the potential for using evolutionary information to track functional variation across the genome, in order to improve whole genome trait prediction. We anticipate that continued work to improve genotype accuracy and deleterious mutation assessment will lead to improved genomic assessments of cassava clones.

plant biology↗

Genome-wide Imputation Using the Practical Haplotype Graph in the Heterozygous Crop Cassava

Genomic applications such as genomic selection and genome-wide association have become increasingly common since the advent of genome sequencing. Genotype imputation makes it possible to infer whole genome information from limited input data, making large sampling for genomic applications more feasible, especially in non-model species where resources are less abundant. Imputation becomes increasingly difficult in heterozygous species where haplotypes must be phased. The Practical Haplotype Graph is a recently developed tool that can accurately impute genotypes, using a reference panel of haplotypes. The Practical Haplotype Graph is a haplotype database that implements a trellis graph to predict haplotypes using minimal input data. Genotyping information is aligned to the database and missing haplotypes are predicted from the most likely path through the graph. We showcase the ability of the Practical Haplotype Graph to impute genomic information in the highly heterozygous crop cassava (Manihot esculenta). Accurately phased haplotypes were sampled from runs of homozygosity across a diverse panel of individuals to populate the graph, which proved more accurate than relying on computational phasing methods. At 1X input sequence coverage, the Practical Haplotype Graph achieves a high concordance between predicted and true genotypes (R=0.84), as compared to the standard imputation tool Beagle (R=0.69). This improved accuracy was especially visible in the prediction of rare and heterozygous alleles. We validate the Practical Haplotype Graph as an accurate imputation tool in the heterozygous crop cassava, showing its potential for application in heterozygous species.

genomics↗