bioRxiv Science⌕ Search

Biology subjects

Endelman, J. B.

Publications and source records attributed to Endelman, J. B..

11 recordsLinked to original sources

Targeted genotyping-by-sequencing of potato and software for imputation

Mid-density targeted genotyping-by-sequencing (GBS) combines trait-specific markers with thousands of genomic markers at an attractive price for linkage mapping and genomic selection. A 2.5K targeted GBS assay for potato was developed using the DArTagTM technology and later expanded to 4K targets. Genomic markers were selected from the potato InfiniumTM SNP array to maximize genome coverage and polymorphism rates. The DArTag and SNP array platforms produced equivalent dendrograms in a test set of 298 tetraploid samples, and 83% of the common markers showed good quantitative agreement, with RMSE (root-mean-squared-error) less than 0.5. DArTag is suited for genomic selection candidates in the clonal evaluation trial, coupled with imputation to a higher density platform for the training population. Using the software polyBreedR, an R package for the manipulation and analysis of polyploid marker data, the RMSE for imputation by linkage analysis was 0.15 in a small half-diallel population (N=85), which was significantly lower than the RMSE of 0.42 with the Random Forest method. Regarding high-value traits, the DArTag markers for resistance to potato virus Y, golden cyst nematode, and potato wart appeared to track their targets successfully, as did multi-allelic markers for maturity and tuber shape. In summary, the potato DArTag assay is a transformative and publicly available technology for potato breeding and genetics. Core IdeasO_LIA mid-density, targeted genotyping-by-sequencing (GBS) assay was developed for potato. C_LIO_LIThe GBS assay includes markers for resistance to potato virus Y, golden cyst nematode, and potato wart. C_LIO_LIThe GBS assay includes multi-allelic markers for potato maturity and tuber shape. C_LIO_LIThe polyBreedR software has functions for manipulating and imputing polyploid marker data in Variant Call Format. C_LIO_LILinkage Analysis was more accurate than the Random Forest method when imputing from 2K to 10K markers. C_LI

genomics↗

Development of KASP markers for the potato virus Y resistance gene Rychc using whole-genome resequencing data

Potato virus Y is the most important potato virus worldwide, affecting tuber yield and quality. The resistance gene Rychc, derived from the potato wild relative Solanum chacoense, provides broad spectrum and durable resistance to the virus and has been used to develop resistant cultivars. Several DNA markers have been developed and have contributed to the efficient selection of resistant individuals. In this study, we developed Kompetitive Allele Specific PCR markers for Rychc using whole-genome resequencing data for a diverse set of 25 PVY susceptible cultivars and a Rychc-positive clone. Marker Ry_4099 targets two variants in the 3-UTR and was able to discriminate all five allele dosages in a tetraploid test population. Marker Ry_3331 targets two variants in Exon 4 and, although it only provides presence/absence information, it discriminates between the two known resistant alleles of Rychc. These markers will greatly contribute to efficient development of resistant cultivars.

plant biology↗

A KASP Marker for the Potato Late Blight Resistance Gene RB/Rpi-blb1

The disease late blight is a threat to potato production worldwide, making genetic resistance an important target for breeding. The resistance gene RB/Rpi-blb1 is effective against most strains of the causal pathogen, Phytophthora infestans. Until now, potato breeders have utilized a Sequence Characterized Amplified Region (SCAR) marker to screen for RB. Our objective was to design and validate a Kompetitive Allele Specific PCR (KASP) marker, which has advantages for high-throughput screening. First, the accuracy of the SCAR marker was confirmed in two segregating tetraploid populations. Then, using whole genome sequencing data for two RB-positive segregants and a diverse set of 23 RB-negative varieties, a SNP in the 5 untranslated (UTR) region was identified as unique to RB. The KASP marker based on this SNP, which had 100% accuracy in the cultivated diversity panel, was used to generate diploid breeding lines containing RB. The KASP marker is publicly available for others to utilize.

genetics↗

Using Haplotype and QTL Analysis to Fix Favorable Alleles in Diploid Potato Breeding

At present, the potato of international commerce is autotetraploid, and the complexity of this genetic system creates limitations for breeding. Diploid potato breeding has long been used for population improvement, and thanks to improved understanding of the genetics of gametophytic self-incompatibility, there is now sustained interest in the development of uniform F1 hybrid varieties based on inbred parents. We report here on the use of haplotype and QTL analysis in a modified backcrossing (BC) scheme, using primary dihaploids of S.tuberosum as the recurrent parental background. In Cycle 1 we selected XD3-36, a self-fertile F2 clone homozygous for the self-compatibility gene Sli. Signatures of gametic and zygotic selection were observed at multiple loci in the F2 generation, including Sli. In the BC1 cycle, an F1 population derived from XD3-36 showed a bimodal response for vine maturity, which led to the identification of late vs. early alleles in XD3-36 for the gene StCDF1 (Cycling DOF Factor 1). Greenhouse phenotypes and haplotype analysis were used to select a vigorous and self-fertile F2 individual with 43% homozygosity, including for Sli and the early-maturing allele StCDF1.3. Partially inbred lines from the BC1 and BC2 cycles have been used to initiate new cycles of selection, with the goal of reaching higher homozygosity while maintaining plant vigor, fertility, and yield. Core IdeasO_LIPartially inbred, diploid potato lines were developed for transitioning to an inbred-hybrid breeding system. C_LIO_LIMulti-generational linkage analysis was used to track and fix favorable alleles without haplotype-specific markers. C_LIO_LISignatures of gametic and zygotic selection were detected by maximum likelihood. C_LI

genetics↗

Fully efficient, two-stage analysis of multi-environment trials with directional dominance and multi-trait genomic selection

Plant breeders interested in genomic selection often face challenges to fully utilizing the multi-trait, multi-environment datasets they rely on for selection. R package StageWise was developed to go beyond the capabilities of most specialized software for genomic prediction, without requiring the programming skills needed for more general-purpose software for mixed models. As the name suggests, one of the core features is a fully efficient, two-stage analysis for multiple environments, in which the full variance-covariance matrix of the Stage 1 genotype means is used in Stage 2. Another feature is directional dominance, including for polyploids, to account for inbreeding depression in outbred crops. StageWise enables selection with multi-trait indices, including restricted indices with one or more traits constrained to have zero response. For a potato dataset with 943 genotypes evaluated over 6 years, including the Stage 1 errors in Stage 2 reduced the Akaike Information Criterion (AIC) by 29, 67, and 104 for maturity, yield, and fry color, respectively. The proportion of variation explained by heterosis was largest for yield but still only 0.03, likely because of limited variation for the genomic inbreeding coefficient. Due to the large additive genetic correlation (0.57) between yield and maturity, naive selection on an index combining yield and fry color led to an undesirable response for later maturity. The restricted index coefficients to maximize genetic merit without delaying maturity were identified. The software and three vignettes are available at https://github.com/jendelman/StageWise.

genetics↗

Clonal breeding strategies to harness heterosis: insights from stochastic simulation

To produce genetic gain, hybrid crop breeding can change the additive as well as dominance genetic value of populations, which can lead to utilization of heterosis. A common hybrid breeding strategy is reciprocal recurrent selection (RRS), in which parents of hybrids are typically recycled within pools based on general combining ability (GCA). However, the relative performance of RRS and other possible breeding strategies have not been thoroughly compared. RRS can have relatively increased costs and longer cycle lengths which reduce genetic gain, but these are sometimes outweighed by its ability to harness heterosis due to dominance and increase genetic gain. Here, we used stochastic simulation to compare gain per unit cost of various clonal breeding strategies with different amounts of population inbreeding depression and heterosis due to dominance, relative cycle lengths, time horizons, estimation methods, selection intensities, and ploidy levels. In diploids with phenotypic selection at high intensity, whether RRS was the optimal breeding strategy depended on the initial population heterosis. However, in diploids with rapid cycling genomic selection at high intensity, RRS was the optimal breeding strategy after 50 years over almost all amounts of initial population heterosis under the study assumptions. RRS required more population heterosis to outperform other strategies as its relative cycle length increased and as selection intensity decreased. Use of diploid fully inbred parents vs. outbred parents with RRS typically did not affect genetic gain. In autopolyploids, RRS typically was not beneficial regardless of the amount of population inbreeding depression. Key MessageReciprocal recurrent selection sometimes increases genetic gain per unit cost in clonal diploids with heterosis due to dominance, but it typically does not benefit autopolyploids.

genetics↗

The genetic architectures of vine and skin maturity in tetraploid potato

Potato vine and skin maturity, which refer to foliar senescence and adherence of the tuber periderm, respectively, are both important to production and therefore breeding. Our objective was to investigate the genetic architectures of these traits in a genome-wide association panel of 586 genotypes, and through joint linkage mapping in a half-diallel subset (N = 397). Skin maturity was measured by image analysis after mechanized harvest 120 days after planting. To correct for the influence of vine maturity on skin maturity under these conditions, the former was used as a covariate in the analysis. The genomic heritability based on a 10K SNP array was 0.33 for skin maturity vs. 0.46 for vine maturity. Only minor QTL were detected for skin maturity, the largest being on chromosome 9 and explaining 8% of the variation. As in many previous studies, S. tuberosum Cycling DOF Factor 1 (CDF1) had a large influence on vine maturity, explaining 33% of the variation in the panel as a bi-allelic SNP and 44% in the half-diallel as a multi-allelic QTL. From the estimated effects of the parental haplotypes in the half-diallel and prior knowledge of the allelic series for CDF1, the CDF1 allele for each haplotype was predicted and ultimately confirmed through whole-genome sequencing. The ability to connect statistical alleles from QTL models with biological alleles based on DNA sequencing represents a new milestone in genomics-assisted breeding for tetraploid species.

genetics↗

QTL Mapping in Outbred Tetraploid (and Diploid) Diallel Populations

Over the last decade, multiparental populations have become a mainstay of genetics research in diploid species. Our goal was to extend this paradigm to autotetraploids by developing software for quantitative trait locus (QTL) mapping in connected F1 populations derived from a set of shared parents. For QTL discovery, phenotypes are regressed on the dosage of parental haplotypes to estimate additive effects. Statistical properties of the model were explored by simulating half-diallel diploid and tetraploid populations with different population sizes and numbers of parents. Across scenarios, the number of progeny per parental haplotype (pph) largely determined the statistical power for QTL detection and accuracy of the estimated haplotype effects. Multi-allelic QTL with heritability 0.2 were detected with 90% probability at 25 pph and genome-wide significance level 0.05, and the additive haplotype effects were estimated with over 90% accuracy. Following QTL discovery, the software enables a comparison of models with multiple QTL and non-additive effects. To illustrate, we analyzed potato tuber shape in a half-diallel population with 3 tetraploid parents. A well-known QTL on chromosome 10 was detected, for which the inclusion of digenic dominance lowered the Deviance Information Criterion (DIC) by 17 points compared to the additive model. The final model also contained a minor QTL on chromosome 1, but higher order dominance and epistatic effects were excluded based on the DIC. In terms of practical impacts, the software is already being used to select offspring based on the effect and dosage of particular haplotypes in breeding programs.

genetics↗

Haplotype Reconstruction in Connected Tetraploid F1 Populations

In diploid species, many multi-parental populations have been developed to increase genetic diversity and quantitative trait loci (QTL) mapping resolution. In these populations, haplotype reconstruction has been used as a standard practice to increase QTL detection power in comparison with the marker-based association analysis. To realize similar benefits in tetraploid species (and eventually higher ploidy levels), a statistical framework for haplotype reconstruction has been developed and implemented in the software PolyOrigin for connected tetraploid F1 populations with shared parents. Haplotype reconstruction proceeds in two steps: first, parental genotypes are phased based on multi-locus linkage analysis; second, genotype probabilities for the parental alleles are inferred in the progeny. PolyOrigin can utilize genetic marker data from single nucleotide polymorphism (SNP) arrays or from sequence-based genotyping; in the latter case, bi-allelic read counts can be used (and are preferred) as input data to minimize the influence of genotype call errors at low depth. To account for errors in the input map, PolyOrigin includes functionality for filtering markers, inferring inter-marker distances, and refining local marker ordering. Simulation studies were used to investigate the effect of several variables on the accuracy of haplotype reconstruction, including the mating design, the number of parents, population size, and sequencing depth. PolyOrigin was further evaluated using an autotetraploid potato dataset with a 3x3 half-diallel mating design. In conclusion, PolyOrigin opens up exciting new possibilities for haplotype analysis in tetraploid breeding populations.

bioinformatics↗

Characterization of a late blight resistance gene homologous to R2 in potato variety Payette Russet

Breeding for late blight resistance has traditionally relied on phenotypic selection, but as the number of characterized resistance (R) genes has grown, so have the possibilities for genotypic selection. One challenge for breeding russet varieties is the lack of information about the genetic basis of resistance in this germplasm group. Based on observations of strong resistance by Payette Russet to genotype US-23 of the late blight pathogen Phytophthora infestans in inoculated experiments, we deduced the variety must contain at least one major R gene. To identify the gene(s), 79 F1 progeny were screened using a detached leaf assay and classified as resistant vs. susceptible. Linkage mapping using markers from the potato SNP array revealed a single resistant haplotype on the short arm of chromosome group 4, which coincides with the R2/Rpi-abpt/Rpi-blb3 locus. PCR amplification and sequencing of the gene in Payette revealed it is homologous to R2, and transient expression experiments in Nicotiana benthamiana confirmed its recognition of the Avr2 effector. Sequencing of a small diversity panel revealed a SNP unique to resistant haplotypes at the R2 locus, which was converted to a KASP marker that showed perfect prediction accuracy in the F1 population and diversity panel. Although many genotypes of P. infestans are virulent against R2, even when defeated this gene may be valuable as one component of a multi-genic approach to quantitative resistance.

genetics↗

Image-based Phenotyping and Genetic Analysis of Potato Skin Set and Color

Image-based phenotyping offers new opportunities for fast, objective, and reliable measurement for breeding and genetics research. In the current study, image analysis was used to quantify potato skin color and skin set, which are critical for the marketability of new varieties. A set of 15 red potato varieties and advanced breeding lines was evaluated over two years at a single location, with two harvest times in the second year. After mechanical harvest and grading, 7-8 representative tubers per plot were photographed, and the photos were analyzed with ImageJ to measure skinning (as % surface area) and skin color using the Hue, Chroma and Lightness (HCL) representation. The plot-based heritability was consistently high (> 0.77) across traits and environments; the genetic correlation between environments was also high, ranging from 0.81 to 0.98. Significant increases in Lightness and Chroma, as well as a decrease in skinning, were observed at the late compared to early harvest, while the opposite trends for color were observed after six weeks of storage. The three color traits were unexpectedly collinear in this study, with the first principal component explaining 86% of the variation. This result may reflect the physiology of red color in potato, but the highly selected nature of the 15 genotypes may also be a factor. Image-based phenotyping offers new opportunities to advance genetic gain and understanding for tuber appearance traits that have been difficult to precisely measure in the past.

genetics↗