bioRxiv Science⌕ Search

Biology subjects

Fuji, Y.

Publications and source records attributed to Fuji, Y..

3 recordsLinked to original sources

Metabolome genome-wide association study reveals hierarchical and epistatic genetic control of flavonoid metabolism in soybean

Metabolic phenotypes are often governed by complex genetic architectures involving both additive and non-additive effects. However, the extent to which epistatic interactions contribute to the pathway-level regulation of plant metabolism remains unclear. In this study, we investigated the genetic architecture of flavonoid-related metabolites using metabolomic and genomic data from 200 soybean accessions cultivated under multiple environmental conditions. Broad-sense heritability estimates revealed that many metabolites were under strong genetic control, particularly flavonoid-related metabolites. Principal component analysis-based metabolome-wide genome-wide association studies identified four major loci associated with flavonoid metabolic variation, including a locus corresponding to flavonoid 3'-hydroxylase. Conditional analyses based on multilocus genetic backgrounds demonstrated that the effects of downstream loci were highly dependent on upstream genotypes. In particular, single-nucleotide polymorphism effects were frequently detectable only in specific allelic backgrounds defined by the major flavonoid 3'-hydroxylase locus, consistent with strong epistatic interactions among loci. Bayesian network analyses further supported a hierarchical genetic structure consistent with upstream regulation of downstream loci across the flavonoid biosynthetic pathway. These results demonstrate that highly heritable metabolic phenotypes can be controlled by a few loci exhibiting both additive and context-dependent non-additive effects. Our findings provide evidence that pathway-level metabolic diversity in soybean is generated through hierarchical and epistatic genetic control involving a limited set of key loci.

plant biology↗

Integration of Proxy Intermediate Omics traits into a Nonlinear Two-Step model for accurate phenotypic prediction

Intermediate omics traits, which mediate the effects of genetic variation on phenotypic traits, are increasingly recognised as valuable components of genetic evaluation. In particular, rhizosphere microbiota play a crucial role in plant health and productivity; however, their complex interactions with host genetics remain challenging to model. Although two-step modeling frameworks have been proposed to integrate intermediate omics traits into phenotype prediction, existing approaches do not incorporate nonlinear relationships between different omics layers. To address this, we have proposed a two-step phenotype prediction framework that integrates genomic, rhizosphere microbiome, and metabolome (meta-metabolome) data, while explicitly capturing omicsomics nonlinearities. The first step is to predict meta-metabolome traits from genetic and microbial features, thus effectively isolating them from the environmental noise. In this process, intermediate "proxy" omics traits are generated as general biological information to provide robust models. The second step utilises this "proxy" to enhance the accuracy of the phenotype prediction. We compared the linear model (Best Linear Unbiased Prediction, BLUP) and the nonlinear model (Random Forest, RF) at each step, as demonstrated through simulations and empirical analysis of a multi-omics soybean dataset in which nonlinear modeling captures intricate omics interactions. Notably, our approach enables phenotype prediction without requiring the original meta-metabolome data used in model training, thereby reducing reliance on costly omics measurements. This framework integrates intermediate omics traits into genomic prediction to improve prediction accuracy and provide solutions for deeper insights into plant-microbiome interactions.

bioinformatics↗

An integrative framework of stochastic variational variable selection for joint analysis of multi-omics microbiome data

High-dimensional multi-omics microbiome data plays an important role in elucidating microbial communities interactions with their hosts and environment in critical diseases and ecological changes. Although Bayesian clustering methods have recently been used for the integrated analysis of multi-omics data, no method designed to analyze multi-omics microbiome data has been proposed. In this study, we propose a novel framework called integrative stochastic variational variable selection (I-SVVS), which is an extension of stochastic variational variable selection for high-dimensional microbiome data. The I-SVVS approach addresses a specific Bayesian mixture model for each type of omics data, such as an infinite Dirichlet multinomial mixture model for microbiome data and an infinite Gaussian mixture model for metabolomic data. This approach is expected to reduce the computational time of the clustering process and improve the accuracy of the clustering results. Additionally, I-SVVS identifies a critical set of representative variables in multi-omics microbiome data. Three datasets from soybean, mice, and humans (each set integrated microbiome and metabolome) were used to demonstrate the potential of I-SVVS. The results indicate that I-SVVS achieved improved accuracy and faster computation compared to existing methods across all test datasets. It effectively identified key microbiome species and metabolites characterizing each cluster. For instance, the computational analysis of soybean dataset, including 377 samples with 16,943 microbiome species and 265 metabolome features, was completed in 2.18 hours using I-SVVS, compared to 2.35 days with Clusternomics and 1.12 days with iClusterPlus. The software for this analysis, written in Python, is freely available at https://github.com/tungtokyo1108/I-SVVS.

bioinformatics↗