bioRxiv Science⌕ Search

Biology subjects

Haley, O. C.

Publications and source records attributed to Haley, O. C..

3 recordsLinked to original sources

Application of RFdiffusion to predict interspecies protein-protein interactions between fungal pathogens and cereal crops

Plant pathogenic fungi secrete small proteins known as effectors which help overcome the plant defense response and cause disease. The concept of effector-triggered immunity in plants evolved from the "gene for gene hypothesis" which describes plant resistance or susceptibility to plant pathogens based on interspecies protein-protein interactions (PPIs) between plant-derived resistance (R) genes and pathogen-derived avirulence (Avr) effector genes. Understanding the molecular interactions mediating these host-pathogen interactions in effector-triggered immunity is thus essential to managing fungal disease. In silico methods of predicting interspecies PPIs have been heavily studied to identify target genes for crop resistance. But conventional sequence-based homology methods (i.e., interlog, domain-based inference) for predicting interspecies PPIs are not as powerful as methods that also incorporate structural homology. The objective of this study was to develop a computational workflow to predict PPIs between pathogenic fungi and their cereal hosts by leveraging recent advances in artificial intelligence and structural biology. This workflow proposes the use of a generative model, RFdiffusion, to predict the structure of truncated segments of proteins likely to bind to query effector proteins. The binder structures were filtered based on the number of contacts at the effectors known binding residues. Acceptable structures were then input into FoldSeek to search the host proteome for host proteins containing similar sub-structures. Experimentally-validated PPIs between rice (Oryzae sativa cv. Japonica) and rice blast fungus (Magnaporthe oryzae) were used for workflow validation. The effects of binder length and the binding residues mode of action (i.e., residues at active/substrate recognition sites) on the binder quality and presumptive host protein matches were explored. Ultimately, 11 out of 14 experimentally validated PPIs were recovered computationally, indicating a high recall (>78%) for the workflow. The shorter binders recovered most of the PPIs, but may have produced the most false positives, as functional analyses revealed that these host proteins displayed a wide variety of functions. These findings emphasize that subject matter expertise is still required to decipher the prediction results. Yet, this framework for elucidating interactions between fungal pathogens and host proteins could provide valuable insight into mechanisms of susceptibility or resistance at a scale friendly to limited computational resources, and facilitate the development of control strategies that reduce crop diseases.

plant biology↗

Fusarium Protein Toolkit: AI-powered tools to combat fungal threats to agriculture

BackgroundThe fungal genus Fusarium poses significant threats to food security and safety worldwide because it consists of numerous species that cause destructive diseases in crops, as well as mycotoxin contamination. The adverse effects of climate change are exacerbating some existing threats and causing new problems. These challenges highlight the need for innovative solutions, including the development of advanced tools to identify targets to control crop diseases and mycotoxin contamination incited by Fusarium. DescriptionIn response to these challenges, we developed the Fusarium Protein Toolkit (FPT, https://fusarium.maizegdb.org/), a web-based tool that allows users to interrogate the structural and variant landscape within the Fusarium pan-genome. FPT offers a comprehensive approach to understanding and mitigating the detrimental effects of Fusarium on agriculture. The tool displays both AlphaFold and ESMFold-generated protein structure models from six Fusarium species. The structures are accessible through a user-friendly web portal and facilitate comparative analysis, functional annotation inference, and identification of related protein structures. Using a protein language model, FPT predicts the impact of over 270 million coding variants in two of the most agriculturally important species, Fusarium graminearum, which causes Fusarium head blight and trichothecene mycotoxin contamination of cereals, and F. verticillioides, which causes ear rot and fumonisin mycotoxin contamination of maize. To facilitate the assessment of naturally occurring genetic variation, FPT provides variant effect scores for proteins in a Fusarium pan-genome constructed from 22 diverse species. The scores indicate potential functional consequences of amino acid substitutions and are displayed as intuitive heatmaps using the PanEffect framework. ConclusionFPT fills a knowledge gap by providing previously unavailable tools to assess structural and missense variation in proteins produced by Fusarium, the most agriculturally important group of mycotoxin-producing plant pathogens. FPT will deepen our understanding of pathogenic mechanisms in Fusarium, and aid the identification of genetic targets that can be used to develop control strategies that reduce crop diseases and mycotoxin contamination. Such targets are vital to solving the agricultural problems incited by Fusarium, particularly evolving threats affected by climate change. By providing a novel approach to interrogate Fusarium-induced crop diseases, FPT is a crucial step toward safeguarding food security and safety worldwide.

bioinformatics↗

PanEffect: A pan-genome visualization tool for variant effects in maize

Understanding the effects of genetic variants is crucial for accurately predicting traits and phenotypic outcomes. Recent advances have utilized protein language models to score all possible missense variant effects at the proteome level for a single genome, but a reliable tool is needed to explore these effects at the pan-genome level. To address this gap, we introduce a new tool called PanEffect. We implemented PanEffect at MaizeGDB to enable a comprehensive examination of the potential effects of coding variants across 51 maize genomes. The tool allows users to visualize over 550 million possible amino acid substitutions in the B73 maize reference genome and also to observe the effects of the 2.3 million natural variations in the maize pan-genome. Each variant effect score, calculated from the Evolutionary Scale Modeling (ESM) protein language model, shows the log-likelihood ratio difference between B73 and all variants in the pan-genome. These scores are shown using heatmaps spanning benign outcomes to strong phenotypic consequences. Additionally, PanEffect displays secondary structures and functional domains along with the variant effects, offering additional functional and structural context. Using PanEffect, researchers now have a platform to explore protein variants and identify genetic targets for crop enhancement. Availability and implementation: The PanEffect code is freely available on GitHub (https://github.com/Maize-Genetics-and-Genomics-Database/PanEffect). A maize implementation of PanEffect and underlying datasets are available at MaizeGDB (https://www.maizegdb.org/effect/maize/).

bioinformatics↗