bioRxiv Science⌕ Search

Biology subjects

Tomura, S.

Publications and source records attributed to Tomura, S..

5 recordsLinked to original sources

Improved Ensemble Performance by Weight Optimisation for the Genomic Prediction of Maize Flowering Time Traits

Ensembles of multiple genomic prediction models have demonstrated improved prediction performance over the individual models contributing to the ensemble. The outperformance of ensemble models is expected from the Diversity Prediction Theorem, which states that for ensembles constructed with diverse prediction models, the ensemble prediction error becomes lower than the mean prediction error of the individual models. While a naive ensemble-average model provides baseline performance improvement by aggregating all individual prediction models with equal weights, optimising weights for each individual model could further enhance ensemble prediction performance. The weights can be optimised based on their level of informativeness regarding prediction error and diversity. Here, we evaluated weighted ensemble-average models with three possible weight optimisation approaches (linear transformation, Nelder-Mead and Bayesian) using flowering time and tillering traits from two maize nested associated mapping (NAM) datasets; TeoNAM and MaizeNAM. The three proposed weighted ensemble-average approaches improved prediction performance in several of the prediction scenarios investigated. In particular, the weighted ensemble models enhanced prediction performance when the adjusted weights differed substantially from the equal weights used by the naive ensemble models. For performance comparisons among the weighted ensembles, there was no clear superiority among the proposed approaches in both prediction accuracy and error across the prediction scenarios. Weight optimisation for ensembles warrants further investigation to explore the opportunities to improve their prediction performance; for example, integration of a weighted ensemble with a simultaneous hyperparameter tuning process may offer a promising direction for further research.

bioinformatics↗

Ensembles of Graph Neural Networks Supervised by Genotype-to-Phenotype Structures Improved Genomic Prediction Performance

Accurate selection of favourable crop genotypes has motivated the exploration of diverse prediction algorithms for crop breeding applications. One genomic prediction method that has not been fully explored is graph attention networks (GAT). By directly analysing graphical data with the attention mechanism, GAT can incorporate the genotype-to-phenotype (G2P) structure to regularise predictions. As one potential G2P structure, a gene network can be inferred from interpretable machine learning models to effectively learn key features of prediction patterns, potentially improving prediction performance. Here, we investigated whether incorporating such data-driven prior knowledge into GAT improved prediction performance compared to GAT models representing a continuum of G2P structures, ranging from infinitesimal to fully connected. Applying the Diversity Prediction Theorem, we also combined these diverse G2P structures into an ensemble of GAT genomic prediction models to integrate complementary strengths of multiple models. The results for flowering time traits in two maize nested association mapping datasets showed a lack of consistent performance improvement in the data-driven prior knowledge GAT model. However, consistent outperformance was observed for the ensemble of GAT models. Improved predictions from the ensemble model may be driven by its ability to capture a more complete representation of the inferred gene network through the integration of information from diverse G2P structures. The observed results using the GAT methodology provided the foundation for potential performance improvement using GAT by integrating biological prior knowledge derived from omics data and empirically verified gene interactions in future research, thereby potentially enhancing the GAT ensemble performance.

genomics↗

Ensemble-based genomic prediction for maize flowering time reveals novel insights into trait genetic architecture and improves prediction for breeding applications

While various genomic prediction models have been evaluated for their potential to accelerate genetic gain for multiple traits, no individual genomic prediction model has outperformed all others across all applications. As an alternative approach, ensembles of multiple individual genomic prediction models can be applied to utilise the complementary strengths of individual prediction models and offset the prediction errors of each. We used the EasiGP (Ensemble AnalySis with Interpretable Genomic Prediction) pipeline to investigate the performance of an ensemble approach, targeting flowering-time traits measured in two maize nested association mapping datasets. For both datasets, the ensemble-based prediction approach achieved higher prediction accuracy and lower prediction error across the flowering-time traits compared to each individual model. Multiple genomic regions known to contain key flowering-time related genes were repeatedly included as features across individual genomic prediction models, indicating the models successfully captured SNPs as features that are associated with genomic regions known to contain flowering-time genes. Although repeatability was high for some genomic regions, estimated marker effects varied across many genomic regions, suggesting that the models might also have captured different aspects of the genetic variation underlying the traits. The ensemble combination of the diverse views likely contributed to the improvement of prediction performance by the ensemble-based approach over the individual prediction models. Ensemble-based prediction can be applied to overcome limitations observed in the continuous exploration for the best individual genomic prediction models that can consistently achieve the highest prediction performance, thereby potentially contributing to improved prediction accuracy for applications in crop breeding. Article summaryThis study targets researchers interested in the performance of genomic prediction models. To demonstrate potential advantages of an ensemble of diverse individual genomic prediction models, we investigated the prediction of key flowering-time traits (days to anthesis and anthesis to silking interval) in two maize datasets. The ensemble approach consistently improved the prediction performance. The improvement was attributed to the offset of prediction errors by combining multiple different dimensions of trait genetic variation. Ensembles can lead to higher selection accuracy of desirable individuals for applications in crop breeding.

bioinformatics↗

Circos Plots for Genome Level Interpretation of Genomic Prediction Models

Ensemble of multiple genomic prediction models have grown in popularity due to consistent prediction performance improvements in crop breeding. However, technical tools that analyse the predictive behaviour at the genome level are lacking. Here, we develop a computational tool called Ensemble AnalySis with Interpretable Genomic Prediction (EasiGP) that uses circos plots to visualise how different genomic prediction models quantify contributions of marker effects to trait phenotypes. As a demonstration of EasiGP, multiple genomic prediction models, spanning conventional statistical and machine learning algorithms, were used to infer the genetic architecture of days to anthesis (DTA) in a maize mapping population. The results indicate that genomic prediction models can capture different views of trait genetic architecture, even when their overall profiles of prediction accuracy are similar. Combinations of diverse views of the genetic architecture for the DTA trait in the TeoNAM study might explain the improved prediction performance achieved by ensembles, aligned with the implication of the Diversity Prediction Theorem. In addition to identifying well-known genomic regions contributing to the genetic architecture of DTA in maize, the ensemble of genomic prediction models highlighted several new genomic regions that have not been previously reported for DTA. Finally, different views of trait genetic architecture were observed across sub-populations, highlighting challenges for between-population genomic prediction. A deeper understanding of genomic prediction models with enhanced interpretability using EasiGP can reveal several critical findings at the genome level from the inferred genetic architecture, providing insights into the improvement of genomic prediction for crop breeding programs. Plain Language SummaryWhile an ensemble of genomic prediction models has been applied in crop breeding, the prediction mechanism has not been well-investigated due to the lack of a computational tool to interpret the predictive behaviour. It is critical to investigate prediction models at the genome level to understand how each model quantifies genomic marker effects contributing to the trait genetic architecture. Hence, we developed a computational tool, the Ensemble AnalySis with Interpretable Genomic Prediction (EasiGP), to investigate the genome features and predictive behaviours of the ensemble. Here, we demonstrate the utility of EasiGP using a maize breeding dataset. EasiGP visualised the genetic architecture from diverse interpretable genomic prediction models and identified several well-known key maize genes. EasiGP also revealed several potential new genomic regions for further investigation. EasiGP helps us investigate the trait genetic architecture that can be utilised to benefit crop breeding. Core ideasO_LIA new computational tool, EasiGP, was created to interpret multiple genomic prediction models at the genomic level C_LIO_LIEasiGP visualises the inferred trait genetic architecture from multiple genomic prediction models with circos plots C_LIO_LIAs a case study, EasiGP highlighted several well-known genes regulating the target trait, days to anthesis C_LIO_LIEasiGP can facilitate the discovery of novel genome regions underlying target traits for further investigation C_LI

bioinformatics↗

Improvements in Prediction Performance of Ensemble Approaches for Genomic Prediction in Crop Breeding

The improvement of selection accuracy of genomic prediction is a key factor in accelerating genetic gain for crop breeding. Traditionally, efforts have focused on developing superior individual genomic prediction models. However, this approach has limitations due to the absence of a consistently "best" individual genomic prediction model, as suggested by the No Free Lunch Theorem. The No Free Lunch Theorem states that the performance of an individual prediction model is expected to be equivalent to the others when averaged across all prediction scenarios. To address this, we explored an alternative method: combining multiple genomic prediction models into an ensemble. The investigation of ensembles of prediction models is motivated by the Diversity Prediction Theorem, which indicates the prediction error of the many-model ensemble should be less than the average error of the individual models due to the diversity of predictions among the individual models. To investigate the implications of the No Free Lunch and Diversity Prediction Theorems, we developed a naive ensemble-average model, which equally weights the predicted phenotypes of individual models. We evaluated this model using two traits influencing crop yield--days to anthesis and tiller number per plant--in the Teosinte Nested Association Mapping dataset. The results show that the ensemble approach increased prediction accuracies and reduced prediction errors over individual genomic prediction models. The advantage of the ensemble was derived from the diverse predictions among the individual models, suggesting the ensemble captures a more comprehensive view of the genomic architecture of these complex traits. These results are in accordance with the expectations of the Diversity Prediction Theorem and suggest that ensemble approaches can enhance genomic prediction performance and accelerate genetic gain in crop breeding programs. Article summaryThis research targets selective breeding industries and researchers developing genomic prediction models to accelerate genetic gain in breeding programs. We applied the concept of an ensemble, combining multiple individual genomic prediction models, to predict key traits (days to anthesis and tiller number per plant) in a crop breeding dataset. Here, we show that an ensemble approach increased prediction accuracies and reduced prediction errors over individual genomic prediction models. These results indicate the potential for ensembles of multiple, diverse genomic prediction models to accelerate genetic gain in breeding programs by increasing the accuracy of selection decisions.

bioinformatics↗