bioRxiv Science⌕ Search

Biology subjects

Garcia-Abadillo, J.

Publications and source records attributed to Garcia-Abadillo, J..

9 recordsLinked to original sources

Genotypic and multi-environment phenotypic evaluation of the lima bean USDA National Plant Germplasm System collection

Lima bean (Phaseolus lunatus L.) is an economically and agronomically important grain legume. Lima beans (or limas) show a range of climatic adaptations with independent domestications in the Andes (large-seeded) and Mesoamerica (small- or medium-seeded). We generated and integrated genotypic and comprehensive field- and laboratory-based phenotypic information for the available accessions in the USDA National Plant Germplasm System collection across multiple environments to inform germplasm utilization in breeding. A total of 810 accessions were genotyped using short-read, low-coverage sequencing. Accession geographic origin and domestication explained population structure. A partially overlapping subset of the panel (n=141-308) was field-evaluated across two years in each of Davis, CA, Central Ferry, WA, and Coachella Valley, CA (the latter was fall-planted for evaluation of photoperiod-sensitive accessions) to assess trait performance in contrasting environments. Agronomic traits such as determinacy and flowering time, and seed traits such as seed coat color and hundred-seed weight, were scored. Macronutrient traits (protein, starch, fat, and ash content) were measured on dry (mature) harvested grain via near-infrared spectroscopy. Genome-wide association analyses identified loci significantly associated with descriptive, agronomic, and seed traits, including orthologs of known genes in common bean and novel candidate regions. Genomic predictive abilities were moderate to high for key traits. Finally, we established a conditional core collection that was constrained to include 211 extensively phenotyped accessions and for which 91 supplemental accessions were selected to maximize genetic diversity from among the genotyped accessions. Overall, these resources provide a foundation to support genomics-assisted breeding of limas.

plant biology↗

Reaction Norm Modeling of High-Dimensional Genomic and Environmental Data Improves Prediction Accuracy in Winter Wheat

Genomic prediction models that account genotype-by-environment (GxE) have the potential to accelerate the rate of genetic gain for yield and agronomic performance, yet relatively few studies have applied GxE prediction in public soft red winter wheat (Triticum aestivum) breeding programs. In this study, we extended a reaction norm-based genomic prediction framework by integrating weather-based environmental covariates to more effectively capture genotype- environment interactions. Key agronomic traits, including seed yield, plant height, test weight, and heading date, were evaluated across 33 environments (location-year) using over 3,200 breeding lines from the North Carolina State University small grains breeding program. Multiple genomic prediction models were compared using several cross-validation (CV) schemes representing common breeding scenarios. Across traits, the reaction norm M5 model, which incorporates both GxE and genotype-by-environmental covariate interactions (GxO), achieved the highest prediction accuracy (PA) in CV2 (predicting incomplete field trials) and CV1 for yield and test weight (predicting new lines). The highest PA was observed for test weight under CV2 (0.54) and for yield under CV1 (0.41). Under CV0 (predicting new environments), the M3 model incorporating GxE produced highest PA across traits, with the greatest accuracy for plant height (0.45), although differences among M2, M3, and M4 were small. Prediction under CV00 (predicting new lines in new environments) remained more challenging, with PA values 0.10 - 0.20 across traits. Overall, our results demonstrate that integrating environmental covariates into genomic prediction models can improve predictive performance across diverse wheat-growing environments in North Carolina, supporting their utility for applied breeding efforts. CORE IDEASO_LIIntegrating genotype-by-environment (GxE) interactions with environmental covariates improves prediction accuracy across environments. C_LIO_LIModel performance varies by prediction scenario, with different approaches performing best for new lines, incomplete trials, or new environments. C_LIO_LIPrediction of new lines in new environments remains challenging. C_LI PLAIN LANGUAGE SUMMARYThis study explores how adding environmental information to genomic prediction models can improve prediction accuracy in a public winter wheat breeding program. Using data from multi-environment trials conducted across diverse conditions in North Carolina, we evaluated statistical models that capture how different wheat lines respond to changing environments. By incorporating weather data, we improved the ability to predict performance across locations and years. These findings provide practical insights for refining selection strategies and accelerating genetic gain in wheat breeding.

genetics↗

Optimizing resource allocation in Miscanthus breeding with sparse testing designs for genomic prediction

Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse biobased products. Increasing biomass yield will increase profitability and environmental benefits, so is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; GxE interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for GxE interaction presented the highest PA and the lowest MSE for CNN (PA: [~]0.77, MSE: [~]0.5) and YDY (PA: [~]0.70, MSE: [~]1.3) while for TCM and AIL these ranged from [~]0.28 to 0.41 and [~]1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.

genomics↗

Multi-trait Multi-environment Genomic Prediction Strategies for Miscanthus sacchariflorus Populations

Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (GxExT) and (2) two single-trait multi-environment (STME) models (with and without GxE interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.

genomics↗

Weather Characterization for Optimizing Genomic Prediction in Miscanthus sacchariflorus

Environmental factors affect crop growth and development thus their consideration across sites and years become essential for genotypic evaluation. Genomic selection (GS) has been broadly implemented to accelerate breeding cycles by skipping field evaluations thus allowing early identification of outperforming genotypes. In this study, 7,740 phenotypic records corresponding to 516 Miscanthus sacchariflorus genotypes evaluated in five locations across three years were considered for analysis. Additionally, environmental data on six weather covariates was implemented to characterize similarities between locations. Different sets of locations of variable sizes were used for model calibration based on two cross-validations (CV00 and CV0) schemes leaving out one location at a time. Predictive ability across locations of the best model varied between 0.45 and 0.90 for both schemes. These results were compared to associate predictive ability in function of weather patterns between training and testing sets to allow models calibration optimization. We found it is feasible to optimize resource allocation by considering environmentally correlated sets. In most cases, the information from only one and, at most, two locations were enough to deliver better results than using all four locations, reducing training sets by up to 75%. The results obtained shed light on helping breeders make informed decisions considering weather data when designing evaluations.

genomics↗

ENHANCING GENOMIC PREDICTION MODELS IN MISCANTHUS POPULATIONS BY INCORPORATING THE GENOTYPE-BY-ENVIRONMENT INTERACTION

Giant Miscanthus giganteus (Mxg) is one of the most promising perennial crops to generate biomass feedstock for bioenergy and biobased products. It is derived from the natural inter-species hybridization of Miscanthus sacchariflorus (Msa) and Miscanthus sinensis (Msi) species, thus population improvement within these species is crucial. Genomic selection (GS) is an attractive option to accelerate breeding of perennial grasses, such as Miscanthus, which requires up to three years of evaluation to produce reliable phenotypic data. Hence, genotypes are observed in multiple years and locations causing inconsistent response patterns from one year to the next, between location, and/or location-by-year combinations. These inconsistencies are known as the genotype-by-environment interaction effect (GxE). Although GS has been successfully implemented in multiple annual crops where straightforward cross-validation schemes exist to assess the levels of predictive ability that can be reached, for perennial crops new cross-validation schemes will help avoid data contamination. Here, we propose a series of cross-validation schemes to evaluate model performance for perennial crops. We perform a case study by analyzing one panel of each species (516 genotypes of Msa, 280 genotypes of Msi) scored for biomass yield at different locations around the world over several years. The results of the different cross-validation schemes provide insights about the usefulness of GS to accelerate the breeding process of Miscanthus species. In addition, leveraging the GxE effects of different types significantly increases predictive ability (up to 10% in Msa and 30% for Msi) compared to the conventional approaches based on main effects only.

genomics↗

Bayesian AMMI-Based Simulation of Genotype x Environment Interactions

Genotype-by-environment interaction (GEI) has been studied to identify environment-stable/favorable genotypes. The GEI simulation could help refine the inference by incorporating tangible factors such as genomic and environmental information. The Bayesian additive main effect and multiplicative interaction (Bayesian AMMI) model captures the genotype-specific responses across environments, reflecting directional relationships between genotypes and environments. Thus, we propose a Bayesian AMMI-based GEI simulation framework that utilizes high-throughput environmental covariance matrices to generate GEI effects with interpretable directional structure. To demonstrate the proposed approach, two simulated phenotypes were assessed under four levels of GEI variance. In the first simulation (Sim1), GEI effects were sampled from a multivariate normal distribution defined by the GEI matrix. In the second simulation (Sim2), GEI effects were generated by extending Sim1 with the Bayesian AMMI model. In both simulations, increasing GEI variance resulted in lower correlations of phenotypes across environments and stronger genotype-specific sensitivity to environmental variation. Across five cross-validation designs, models accounting for GEI consistently outperformed one that did not, with prediction accuracy generally decreasing as GEI variance increased. Clear distinctions between the two simulated phenotypes were evident from biplot analyses: Sim2 successfully captured environmental relatedness and genotype-specific responses, whereas such structure was absent in Sim1. These results demonstrate that the proposed Bayesian AMMI-based GEI simulation framework enables interpretable visualization of GEI and supports genomic selection strategies under complex environmental conditions.

bioinformatics↗

Sparse Testing Designs for Optimizing Predictive Ability in Sugarcane Populations

Sugarcane is a crucial crop for sugar and bioenergy production. Saccharose content and total weight are the two main key commercial traits that compose sugarcanes yield. These traits are under complex genetic control and their response patterns are influenced by the genotype-by-environment (GxE) interaction. An efficient breeding of sugarcane demands an accurate assessment of the genotype stability through multi-environment trials (METs), where genotypes are tested/evaluated across different environments. However, phenotyping all genotype-in-environment combinations is often impractical due to cost and limited availability of propagation-materials. This study introduces the sparse testing designs as a viable alternative, leveraging genomic information to predict unobserved combinations through genomic prediction models. This approach was applied to a dataset comprising 186 genotypes across six environments (6 x 186 = 1,116 phenotypes). Our study employed three predictive models, including environment and genotype as main effects, as well as the GxE interaction to predict saccharose accumulation (SA) and tons of cane per hectare (TCH). Calibration sets sizes varying between 72 (6.5%) to 186 (16.7%) of the total number of phenotypes were composed to predict the remaining 930 (83.3%). Additionally, we explored the optimal number of common genotypes across environments for GxE pattern prediction. Results demonstrate that maximum accuracy for SA ({rho} = 0.611) and for TCH ({rho} = 0.341) was achieved using in training sets few (3) to no common (0) genotype across environments maximizing the number of different genotypes that were tested only once. Significantly, we show that reducing phenotypic records for model calibration has minimal impact on predictive ability, with sets of 12 non-overlapped genotypes per environment (72 = 12 x 6) being the most convenient cost-benefit combination.

genomics↗

Introducing CHiDO a No Code Genomic Prediction Software implementation for the Characterization & Integration of Driven Omics

Climate change represents a significant challenge to global food security by altering environmental conditions critical to crop growth. Plant breeders can play a key role in mitigating these challenges by developing more resilient crop varieties; however, these efforts require significant investments in resources and time. In response, it is imperative to use current technologies that assimilate large biological and environmental datasets into predictive models to accelerate the research, development, and release of new improved varieties. Leveraging large and diverse data sets can improve the characterization of phenotypic responses due to environmental stimuli and genomic pulses. A better characterization of these signals holds the potential to enhance our ability to predict trait performance under changes in weather and/or soil conditions with high precision. This paper introduces CHiDO, an easy-to-use, no-code platform designed to integrate diverse omics datasets and effectively model their interactions. With its flexibility to integrate and process data sets, CHiDOs intuitive interface allows users to explore historical data, formulate hypotheses, and optimize data collection strategies for future scenarios. The platforms mission emphasizes global accessibility, democratizing statistical solutions for situations where professional ability in data processing and data analysis is not available. Core ideasO_LIThe authors developed CHiDO, a platform for breeders to build predictive models integrating multi-omics data. C_LIO_LICHiDO is a no-code tool that leverages the reaction norm model proposed by Jarquin et al. (2014). C_LIO_LIThe platform aims to increase access to predictive analytics lowering relevant technical and financial barriers. C_LI

genomics↗