bioRxiv Science⌕ Search

Biology subjects

Harper, L. C.

Publications and source records attributed to Harper, L. C..

2 recordsLinked to original sources

FASSO: An AlphaFold based method to assign functional annotations by combining sequence and structure orthology

Methods to predict orthology play an important role in bioinformatics for phylogenetic analysis by identifying orthologs within or across any level of biological classification. Sequence-based reciprocal best hit approaches are commonly used in functional annotation since orthologous genes are expected to share functions. The process is limited as it relies solely on sequence data and does not consider structural information and its role in function. Previously, determining protein structure was highly time-consuming, inaccurate, and limited to the size of the protein, all of which resulted in a structural biology bottleneck. With the release of AlphaFold, there are now over 200 million predicted protein structures, including full proteomes for dozens of key organisms. The reciprocal best structural hit approach uses protein structure alignments to identify structural orthologs. We propose combining both sequence- and structure-based reciprocal best hit approaches to obtain a more accurate and complete set of orthologs across diverse species, called Functional Annotations using Sequence and Structure Orthology (FASSO). Using FASSO, we annotated orthologs between five plant species (maize, sorghum, rice, soybean, Arabidopsis) and three distance outgroups (human, budding yeast, and fission yeast). We inferred over 270,000 functional annotations across the eight proteomes including annotations for over 5,600 uncharacterized proteins. FASSO provides confidence labels on ortholog predictions and flags potential misannotations in existing proteomes. We further demonstrate the utility of the approach by exploring the annotation of the maize proteome.

bioinformatics↗

An updated nomenclature for plant ribosomal protein genes

Ban et al. (2014) proposed a nomenclature for ribosomal proteins (r-proteins) that reflects the current understanding of ribosomal protein evolution. In the past few years, this nomenclature has been widely adopted among biomedical researchers and microbiologists. This homology-based r-protein nomenclature has not been as widely adopted among plant biologists, however, presumably because r-protein nomenclature is much more complicated in plants due to gene duplication. Here, we propose compatible upgrades to the homology-guided nomenclature proposed by Ban et al. (2014) so that this naming system can be adopted for widespread use in the plant biology community. We note that Lan et al. (2022) recently proposed updated nomenclature for plant cytosolic ribosomal proteins, focused on Arabidopsis and rice. The nomenclature outlined here is an extension of that proposed by Lan et al. (2022), expanding to include organellar ribosomes and additional species, with the intent that this nomenclature can serve as a template to guide future plant genome annotations. A more detailed comparison highlighting how this naming system builds on the Ban et al. (2014) and Lan et al. (2022) nomenclatures is offered below. At this time, we request community feedback on this proposed nomenclature so that the naming system ultimately chosen represents a broad consensus. Feedback can be communicated to the this working group at plantribosome@gmail.com before July 25th, 2022. Coauthors of this letter and anyone in the scientific community expressing significant interest will then discuss this feedback as a group, reach a consensus agreement, and communicate the updated nomenclature rules through a letter to the editor (expected to be published at The Plant Cell) and the databases at TAIR and MaizeGDB.

plant biology↗