bioRxiv ScienceSearch

Biology subjects

Morota, G.

Publications and source records attributed to Morota, G..

7 recordsLinked to original sources

Leveraging breeding values obtained from random regression models for genetic inference of longitudinal traits

Understanding the genetic basis of dynamic plant phenotypes has largely been limited due to lack of space and labor resources needed to record dynamic traits, often destructively, for a large number of genotypes. However, the recent advent of image-based phenotyping platforms has provided the plant science community with an effective means to non-destructively evaluate morphological, developmental, and physiological processes at regular, frequent intervals for a large number of plants throughout development. The statistical frameworks typically used for genetic analyses (e.g. genome-wide association mapping, linkage mapping, and genomic prediction) in plant breeding and genetics are not particularly amenable for repeated measurements. Random regression (RR) models are routinely used in animal breeding for the genetic analysis of longitudinal traits, and provide a robust framework for modeling traits trajectories and performing genetic analysis simultaneously. We recently used a RR approach for genomic prediction of shoot growth trajectories in rice using 33,674 SNPs. In this study, we have extended this approach for genetic inference by leveraging genomic breeding values derived from RR models for rice shoot growth during early vegetative development. This approach provides improvements over a conventional single time point analyses for discovering loci associated with shoot growth trajectories. The RR approach uncovers persistent, as well as time-specific, transient quantitative trait loci. This methodology can be widely applied to understand the genetic architecture of other complex polygenic traits with repeated measurements. O_LSTCore Ideas:C_LSTO_LIRandom regression models are an appealing framework for GWAS of longitudinal traits C_LIO_LIThis approach provides improvements over a conventional single time point analyses for GWAS C_LIO_LIWe identify QTL with transient and persistent effects on shoot growth in rice C_LI

genetics

Genomic Bayesian confirmatory factor analysis and Bayesian network to characterize a wide spectrum of rice phenotypes

Drawing biological inferences from large data generated to dissect the genetic basis of complex traits remains a challenge. Since multiple phenotypes likely share mutual relationships, elucidating the interdependencies among economically important traits can accelerate the genetic improvement of plants and animals. A Bayesian network depicts a probabilistic directed acyclic graph representing conditional dependencies among variables. This study aims to characterize various phenotypes in rice (Oryza sativa) via confirmatory factor analysis and Bayesian network. Confirmatory factor analysis under the Bayesian treatment hypothesized that 48 observed phenotypes resulted from six latent variables including grain morphology, morphology, flowering time, physiology (e.g., ion content), yield, and morphological salt response. This was followed by studying the genetics of each latent variable. Bayesian network structures involving the genomic component of six latent variables were established by fitting four different algorithms. Negative genomic correlations were obtained between salt response and yield, salt response and grain morphology, salt response and physiology, and morphology and yield, whereas a positive correlation was obtained between yield and grain morphology. There were four common directed edges across the different Bayesian networks. Physiological components influenced the flowering time and grain morphology, and morphology and 4 grain morphology influenced yield. This work suggests that the Bayesian network coupled with factor analysis can provide an effective approach to understand the interdependence patterns among phenotypes and to predict the potential influence of external interventions or selection related to target traits in the high-dimensional interrelated complex traits systems.

genetics

ShinyAIM: Shiny-based Application of Interactive Manhattan Plots for Longitudinal Genome-Wide Association Studies

Due to advancements in sensor-based, non-destructive phenotyping platforms, researchers are increasingly collecting data with higher temporal resolution. These phenotypes collected over several time points are cataloged as longitudinal traits and used for genome-wide association studies (GWAS). Longitudinal GWAS typically yield a large number of output files, posing a significant challenge for data interpretation and visualization. Efficient, dynamic, and integrative data visualization tools are essential for the interpretation of longitudinal GWAS results for biologists but are not widely available to the community. We have developed a flexible and user-friendly Shiny-based online application, ShinyAIM, to dynamically view and interpret temporal GWAS results. The main features of the application include (i) an interactive Manhattan plots for single time points, (ii) a grid plot to view Manhattan plots for all time points simultaneously, (iii) dynamic scatter plots for p-value-filtered selected markers to investigate co-localized genomic regions across time points, (iv) and interactive phenotypic data visualization to capture variation and trends in phenotypes. The application is written entirely in the R language and can be used with limited programming experience. ShinyAIM is deployed online as a Shiny web server application at https://chikudaisei.shinyapps.io/shinyaim/, enabling easy access for users without installation. The application can also be launched on the local machine in RStudio.

genetics

Utilizing random regression models for genomic prediction of a longitudinal trait derived from high-throughput phenotyping

The accessibility of high-throughput phenotyping platforms in both the greenhouse and field, as well as the relatively low cost of unmanned aerial vehicles, have provided researchers with an effective means to characterize large populations throughout the growing season. These longitudinal phenotypes can provide important insight into plant development and responses to the environment. Despite the growing use of these new phenotyping approaches in plant breeding, the use of genomic prediction models for longitudinal phenotypes is limited in major crop species. The objective of this study is to demonstrate the utility of random regression (RR) models using Legendre polynomials for genomic prediction of shoot growth trajectories in rice (Oryza sativa). An estimate of shoot biomass, projected shoot area (PSA), was recored over a period of 20 days for a panel of 357 diverse rice accessions using an image-based greenhouse phenotyping platform. A RR that included a fixed second-order Legendre polynomial, a random second-order Legendre polynomial for the additive genetic effect, a first-order Legendre polynomial for the environmental effect, and heterogeneous residual variances was used to model PSA trajectories. The utility of the RR model over a single time point (TP) approach, where PSA is fit at each time point independently, is shown through four prediction scenarios. In the first scenario, the RR and TP approaches were used to predict PSA for a set of lines lacking phenotypic data. The RR approach showed a 11.6% increase in prediction accuracy over the TP approach. Much of this improvement could be attributed to the greater additive genetic variance captured by the RR approach. The remaining scenarios focused forecasting future phenotypes using a subset of early time points for known lines with phenotypic data, as well new lines lacking phenotypic data. In all cases, PSA could be predicted with high accuracy (r: 0.79 to 0.89 and 0.55 to 0.58 for known and unknown lines, respectively). This study provides the first application of RR models for genomic prediction of a longitudinal trait in rice, and demonstrates that RR models can be effectively used to improve the accuracy of genomic prediction for complex traits compared to a TP approach.

genomics

Including phenotypic causal networks in genome-wide association studies using mixed effects structural equation models

BackgroundPhenotypic networks describing putative causal relationships among multiple phenotypes can be used to infer single-nucleotide polymorphism (SNP) effects in genome-wide association studies (GWAS). In GWAS with multiple phenotypes, reconstructing underlying causal structures among traits and SNPs using a single statistical framework is essential for understanding the entirety of genotype-phenotype maps. A structural equation model (SEM) can be used for such purposes.\n\nMethodsWe applied SEM to GWAS (SEM-GWAS) in chickens, taking into account putative causal relationships among body weight (BW), breast meat (BM), hen-house production (HHP), and SNPs. We assessed the performance of SEM-GWAS by comparing the model results with those obtained from traditional multi-trait association analyses (MTM-GWAS).\n\nResultsThree different putative causal path diagrams were inferred from highest posterior density (HPD) intervals of 0.75, 0.85, and 0.95 using the inductive causation algorithm. A positive path coefficient was estimated for BM[->]BW, and negative values were obtained for BM[->]HHP and BW[->]HHP in all implemented scenarios. Further, the application of SEM-GWAS enabled the decomposition of SNP effects into direct, indirect, and total effects, identifying whether a SNP effect is acting directly or indirectly on a given trait. In contrast, MTM-GWAS only captured overall genetic effects on traits, which is equivalent to combining the direct and indirect SNP effects from SEMGWAS.\n\nConclusionsAlthough MTM-GWAS and SEM-GWAS use the same probabilistic models, we provide evidence that SEM-GWAS captures complex relationships and delivers a more comprehensive understanding of SNP effects compared to MTM-GWAS. Our results showed that SEM-GWAS provides important insight regarding the mechanism by which identified SNPs control traits by partitioning them into direct, indirect, and total SNP effects.

genetics

ShinyGPAS: Interactive genomic prediction accuracy simulator based on deterministic formulas

BackgroundDeterministic formulas highlight the relationships among prediction accuracy and potential factors influencing prediction accuracy prior to performing computationally intensive cross-validation. Visualizing such deterministic formulas in an interactive manner may lead to a better understanding of how genetic factors control prediction accuracy.\n\nResultsThe software to simulate deterministic formulas for genomic prediction accuracy was implemented in R and encapsulated as a web-based Shiny application. ShinyGPAS (Shiny Genomic Prediction Accuracy Simulator) simulates various deterministic formulas and delivers dynamic scatter plots of prediction accuracy vs. genetic factors impacting prediction accuracy, while requiring only mouse navigation in a web browser. ShinyGPAS is available at: https://chikudaisei.shinyapps.io/shinygpas/.\n\nConclusionShinyGPAS is a shiny-based interactive genomic prediction accuracy simulator using deterministic formulas. It can be used for interactively exploring potential factors influencing prediction accuracy in genome-enabled prediction, simulating achievable prediction accuracy prior to genotyping individuals, or supporting in-class teaching. ShinyGPAS is open source software and it is hosted online as a freely available web-based resource with an intuitive graphical user interface.

genetics

Genomic Relatedness Strengthens Genetic Connectedness Across Management Units

Genetic connectedness refers to a measure of genetic relatedness across management units (e.g., herds and flocks). With the presence of high genetic connectedness in management units, best linear unbiased prediction (BLUP) is known to provide reliable comparisons between genetic values. Genetic connectedness has been studied for pedigree-based BLUP; however, relatively little attention has been paid to using genomic information to measure connectedness. In this study, we assessed genome-based connectedness across management units by applying prediction error variance of difference (PEVD), coefficient of determination (CD), and prediction error correlation (r) to a combination of computer simulation and real data (mice and cattle). We found that genomic information (G) increased the estimate of connectedness among individuals from different management units compared to that based on pedigree (A). A disconnected design benefited the most. In both datasets, PEVD and CD statistics inferred increased connectedness across units when using G- rather than A-based relatedness suggesting stronger connectedness. With r once using allele frequencies equal to one-half or scaling G to values between 0 and 2, which is intrinsic to A, connectedness also increased with genomic information. However, PEVD occasionally increased, and r decreased when obtained using the alternative form of G, instead suggesting less connectedness. Such inconsistencies were not found with CD. We contend that genomic relatedness strengthens measures of genetic connectedness across units and has the potential to aid genomic evaluation of livestock species.\n\nThe problem of connectedness or disconnectedness is particularly important in genetic evaluation of managed populations such as domesticated livestock. When selecting among animals from different management units (e.g., herds and flocks), caution is needed; choosing one animal over others across management units may be associated with greater uncertainty than selection within management units. Such uncertainty is reduced if individuals from different management units are genetically linked or connected. In such a case, best linear unbiased prediction (BLUP) offers meaningful comparison of the breeding values across management units for genetic evaluation (e.g., Kuehn et al., 2007).\n\nStructures of breeding programs have a direct influence on levels of connectedness. Wide use of artificial insemination (AI) programs generally increases genetic connectedness across management units. For example, dairy cattle populations are considered highly connected due to dissemination of genetic material from a small number of highly selected sires. The situation may be different for species with less use of AI and more use of natural service mating such as for beef cattle or sheep populations. Under these scenarios, the magnitude of connectedness across management units is reduced and genetic links are largely confined within management units.\n\nPedigree-based genetic connectedness has been evaluated and applied in practice (e.g., Kuehn et al., 2009; Eikje and Lewis, 2015). However, there is a relative paucity of use of genomic information such as single nucletide polymorphisms (SNPs) to ascertain connectedness. It still remains elusive in what scenarios genomics can strengthen connectedness and how much gain can be expected relative to use of pedigree information alone. Connectedness statistics have been used to optimize selective genotyping and phenotyping in simulated livestock (Pszczola et al., 2012) and plant populations (Maenhout et al., 2010), and in real maize (Rincent et al., 2012; Isidro et al., 2015), and real rice data (Isidro et al., 2015). These studies concluded that the greater the connectedness between the reference and validation populations, the greater the predictive performance. However, 1) connectedness among different management units and 2) differences in connectedness measures between pedigree and genomic relatedness were not explored in those studies. For better understanding of genome-based connectedness, it is critical to examine how the presence of management units comes into play. For instance, genomic relatedness provides relationships between distant individuals that appear disconnected according to the pedigree information. In addition, it captures Mendelian sampling that is not present in pedigree relationships (Hill and Weir, 2011). Thus, genomic information is expected to strengthen measures of connectedness, which in turn refines comparisons of genetic values across different management units. The objective of this study was to assess measures of genetic connectedness across management units with use of genomic information. We leveraged the combination of real data and computer simulation to compare gains in measures of connectedness when moving from pedigree to genomic relationships. First, we studied a heterogenous mice dataset stratified by cage. Then we investigated approaches to measure connectedness using real cattle data coupled with simulated management units to have greater control over the degree of confounding between fixed management groups and genetic relationships.

genetics