Pangenomics and machine learning reveal genetic variation to optimize carotenoids in sorghum grain
Enhancing grain carotenoid concentration to increase yellowness and provitamin A content in sorghum (Sorghum bicolor [L.] Moench) is a major breeding objective across sub-Saharan Africa. Carotenoid breeding is constrained by costly post-harvest phenotyping and can be accelerated through genetic markers that capture functional variation in key carotenoid biosynthesis genes. Zeaxanthin epoxidase (ZEP) is a major gene controlling sorghum grain carotenoid accumulation, and SNP-based KASP markers facilitate selection of high-carotenoid grain, although quantitative variation remains among lines carrying the favorable allele. We hypothesized that additional functional variation within ZEP and other carotenoid biosynthesis pathway genes contributes to this variation and can be revealed by a pangenome-informed analysis. Using a 33-member pangenome reference, we characterized sequence and structural variation at ZEP and other key carotenoid biosynthesis genes and applied pangenome-derived genotyping to study marker diversity associated with carotenoid accumulation. Evaluation of ZEP in the pangenome reference revealed previously uncharacterized structural variation that is absent from the BTx623 primary reference genome. A pangenome-based association analysis identified genetic markers associated with the accumulation of carotenoids which were specific to particular pangenome reference members, including in RTx430 and SRN39, both yellow endosperm lines. Machine learning identified markers in ZEP, {beta}-OH, ZDS, and Z-ISO as the most predictive for all carotenoid traits, suggesting a multi-locus genetic architecture underlies the accumulation of carotenoids in sorghum grain. By integrating sorghum pangenomic resources with machine learning, this study establishes a framework for pangenome-accelerated trait discovery and identifies new genetic targets for carotenoid biofortification in sorghum.