bioRxiv Science⌕ Search

Biology subjects

Turkus, J.

Publications and source records attributed to Turkus, J..

8 recordsLinked to original sources

Embeddings from standardized sorghum leaf images capture variation in disease response that human scoring misses

Ordinal scoring of plant disease severity by human raters compresses variation in lesion color, size, and number. Inter-rater variability further complicates comparisons and integrated analyses across environments. We developed a low-cost portable imaging chamber to rapidly image large numbers of leaves under standardized lighting, orientation, and backdrops in the field and employed this system to image more than 11,000 leaves across three states. Embeddings from vision encoders predicted human-assigned disease severity scores. No significant GWAS hits were identified using human-assigned or vegetation-index-based disease severity scores, but GWAS using embeddings identified twelve genomic hotspots controlling leaf appearance. Nine were linked to variation in disease symptom severity. Five hotspots corresponded to previously characterized sorghum genes: all three hotspots not linked to disease and two of the nine that were. Roughly one-third of tested embedding--hotspot associations replicated across at least two states, and twenty replicated across all three. eQTL, PheWAS, and large-effect variant analyses identified single candidate genes with plausible mechanistic links to disease symptom severity for six of the seven hotspots not mapping to characterized genes. These results demonstrate the power of combining scalable, standardized leaf imaging with pretrained image encoders to capture genetically controlled variation in diverse disease symptoms that human ordinal scoring misses.

plant biology↗

Predicting complex phenotypes using multi-omics data in maize

Understanding and predicting complex traits in plants remains a fundamental challenge due to the emergent nature of most phenotypes and their dependence on genetic, regulatory, and environmental interactions. Accurate prediction of traits and identification of underlying genetic elements has broad applications for plant breeding, systems biology, and biotechnology. Here, we tested if multi-omic datasets could improve predictive accuracy of 129 diverse maize phenotypes across nine environments using genomic markers, field based transcriptomic data from two locations, and drone-derived phenomic data of vegetative indices. We trained and compared linear (rrBLUP) and nonlinear (support vector regression) models using single- and multi-omics inputs. Multi-omics models consistently outperformed single-omics models for most traits, with genomic and transcriptomic inputs contributing distinct biological features. Phenomic features alone yielded the lowest predictive power but improved predictions for specific trait categories like root architecture. Transcriptomic datasets enabled cross-environment prediction, demonstrating that gene expression patterns from one field site could accurately predict traits measured in another. Environment-specific expression of benchmark flowering time genes highlighted the value of transcriptomics in capturing genotype-by-environment (GxE) interactions not detectable through genomic data alone. These findings demonstrate that integrating transcriptomic and phenomic data with genotypes enhances trait prediction, improves model generalizability across environments, and provides deeper insight into the genetic and regulatory architecture of agriculturally important traits in maize.

plant biology↗

Assessing the impact of yield plasticity on hybrid performance in maize

Improving crop resilience in the face of increasingly extreme and unpredictable weather and reduced access to agricultural inputs such as nitrogen fertilizer and water will require an improved understanding of phenotypic plasticity in crops. To understand the roles of different component traits in determining overall plasticity for grain yield, we generated data from a panel of 122 maize (Zea mays) hybrids grown in replicated field trials in 34 environments spanning 700 miles (1126 km) of the U.S. Corn Belt. We observed that the levels of genetic versus environmental control and the relationships between mean parent release year, overall performance, and linear plasticity were trait-dependent across the 18 agronomic and yield components studied. Importantly and unexpectedly, we observed no clear tradeoff between linear plasticity and mean performance and found only rare examples where genotype-by-environment interactions would alter selection decisions based on the environments tested in our dataset. Furthermore, we showed that overall plasticity was repeatable and appears to be under considerable genetic control but that plasticity in response to nitrogen fertilization was not, which may help explain the limited success in breeding for nitrogen use efficiency. Together, these findings improve our understanding of phenotypic plasticity, with implications for maize breeding.

plant biology↗

Genes and pathways determining flowering time variation in temperate adapted sorghum

The timing of the transition from vegetative to reproductive growth is determined by a complex genetic architecture integrating signals from a diverse set of external and internal stimuli and plays a key role in determining plant fitness and adaptation. However, significant divergence in the identities and functions of many flowering time pathway components has been reported among plant species. Here we employ a combination of genome and transcriptome wide association studies to identify genetic determinants of variation in flowering time across multiple environments in a large panel of primarily photoperiod-insensitive sorghum (Sorghum bicolor), a major crop that has, to date, been the subject of substantially less genetic investigation than its relatives. Gene families that form core components of the flowering time pathway in other species, FT-like and SOC1-like genes, appear to play similar roles in sorghum, but the genes identified are not orthologous to the primary FT-like or SOC1-like genes which play similar roles in related species. The aging pathway appears to play a significant role in determining non-photoperiod determined variation in flowering time in sorghum. Two components of this pathway were identified in a transcriptome wide association study, while a third was identified via genome wide association. Our results demonstrate that, while the functions of larger gene families are conserved, functional data from even closely related species is not a reliable guide to which gene copies will play roles in determining natural variation in flowering time.

plant biology↗

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcripts impacts on plant metabolism has employed data from controlled environments. Here we generate and employ data on variation in the abundance of twenty-six metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies) and transcripts (through both transcriptome-wide association studies and an explainable AI approach based on the random forest) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

genomics↗

Population level gene expression can repeatedly link genes to functions in maize

Transcriptome-Wide Association Studies (TWAS) can provide single gene resolution for candidate genes in plants, complementing Genome-Wide Association Studies (GWAS) but efforts in plants have been met with, at best, mixed success. We generated expression data from 693 maize genotypes, measured in a common field experiment, sampled over a two-hour period to minimize diurnal and environmental effects, using full-length RNA-seq to maximize the accurate estimation of transcript abundance. TWAS could identify roughly ten times as many genes likely to play a role in flowering time regulation as GWAS conducted data from the same experiment. TWAS using mature leaf tissue identified known true positive flowering time genes known to act in the shoot apical meristem, and trait data from new environments enabled the identification of additional flowering time genes without the need for new expression data. eQTL analysis of TWAS-tagged genes identified at least one additional known maize flowering time gene through trans-eQTL interactions. Collectively these results suggest the gene expression resource described here can link genes to functions across different plant phenotypes expressed in a range of tissues and scored in different experiments.

plant biology↗

A Common Resequencing-Based Genetic Marker Dataset for Global Maize Diversity

Maize (Zea mays ssp. mays) populations exhibit vast amounts of genetic and phenotypic diversity. As sequencing costs have declined, an increasing number of projects have sought to measure genetic differences between and within maize populations using whole genome resequencing strategies, identifying millions of segregating single-nucleotide polymorphisms (SNPs) and insertions/deletions (InDels). Unlike older genotyping strategies like microarrays and genotyping by sequencing, resequencing should, in principle, frequently identify and score common genetic variants. However, in practice, different projects frequently employ different analytical pipelines, often employ different reference genome assemblies, and consistently filter for minor allele frequency within the study population. This constrains the potential to reuse and remix data on genetic diversity generated from different projects to address new biological questions in new ways. Here we employ resequencing data from 1,276 previously published maize samples and 239 newly resequenced maize samples to generate a single unified marker set of [~]366 million segregating variants and [~]46 million high confidence variants scored across crop wild relatives, landraces as well as tropical and temperate lines from different breeding eras. We demonstrate that the new variant set provides increased power to identify known causal flowering time genes using previously published trait datasets, as well as the potential to track changes in the frequency of functionally distinct alleles across the global distribution of modern maize.

plant biology↗

Development of the Wheat Practical Haplotype Graph Database as a Resource for Genotyping Data Storage and Genotype Imputation

To improve the efficiency of high-density genotype data storage and imputation in bread wheat (Triticum aestivum L.), we applied the Practical Haplotype Graph (PHG) tool. The wheat PHG database was built using whole-exome capture sequencing data from a diverse set of 65 wheat accessions. Population haplotypes were inferred for the reference genome intervals defined by the boundaries of the high-quality gene models. Missing genotypes in the inference panels, composed of wheat cultivars or recombinant inbred lines genotyped by exome capture, genotyping-by-sequencing (GBS), or whole-genome skim-seq sequencing approaches, were imputed using the wheat PHG database. Though imputation accuracy varied depending on the method of sequencing and coverage depth, we found 93% imputation accuracy with 0.01x sequence coverage, which was only slightly lower than the accuracy obtained using the 0.5x sequence coverage (96.9%). Compared to Beagle, on average, PHG imputation was ~4% (p-value = 0.00027) more accurate, and showed 27% higher accuracy at imputing a rare haplotype introgressed from a wild relative into wheat. The reduced accuracy of imputation with GBS data (90.4%) is likely associated with the small overlap between GBS markers and the exome capture dataset, which was used for constructing PHG. The highest imputation accuracy was obtained with exome capture for the wheat D genome, which also showed the highest levels of linkage disequlibrium and proportion of identity-by-descent regions among accessions in our reference panel. We demonstrate that genetic mapping based on genotypes imputed using PHG identifies SNPs with a broader range of effect sizes that together explain a higher proportion of genetic variance for heading date and meiotic crossover rate compared to previous studies.

genomics↗