bioRxiv Science⌕ Search

Biology subjects

Atencio, H. M.

Publications and source records attributed to Atencio, H. M..

2 recordsLinked to original sources

Comprehensive characterization of the complex BAHD acyltransferase family from 218 land plants species: phylogenomic analysis and identification of specificity determinant positions

Chemodiversity is a fundamental trait acquired by plants during their lands colonization. This resulted from an evolutionary process leading to the increase in the number of homologues from a distinct set of protein superfamilies, many of them associated to the specialized metabolism, which allowed the expansion of the chemical space to cope with several environmental cues. BAHD acyltransferases are among these important superfamilies, catalyzing a reaction leading to the acylation of acceptor metabolites with Coenzyme A-activated donors. BAHD acyltransferases can use a wide variety of substrates and they often times display substrate permissiveness towards a wide variety of substrates. Together, these factors complicates the reliable identification and functional annotation of BAHD homologues, also due to the (relatively) limited amount of biochemical data on BAHD acyltransferases. In this work, we take a phylogenomics and computational approach to study the BAHD superfamily in land plants. Using a clustered training set with 27 proteomes, followed by the classification of additional 191 proteomes, we obtained a final BAHDome with 15607 homologues. The training set was clustered in 16 groups, that, together with the identification of cluster of orthologues from the complete BAHDome, were partially assigned functionally to different families guided by the characterized activities (in terms of metabolites used) present in each group. However, the function assignation was not direct in several cases, due to large sequence number, taxonomical distribution on the group, and high sequence variability intra-group. Finally, we used the different families (and functional subfamilies) identified to detect specificity determining positions (SDPs), that may account to explain the functional diversity by finding key positions in, for example, substrate interaction. However, only a handful of the SDPs identified in this work are linked to the substrate binding pocket. Together, these results allow to partially annotate functional BAHD families, which have evolved in a complex pattern of taxonomical and functional signals to allow interaction with multiple substrates, more likely associated to protein dynamics rather than direct substrate interaction.

genomics↗

Seqrutinator: Non-Functional Homologue Sequence Scrutiny for the Generation of large Datatsets for Protein Superfamily Analysis

BackgroundIn recent years protein bioinformatics has resulted in many good algorithms for multiple sequence alignment (MSA) and phylogeny. Little attention has been paid to sequence selection whereas notably recently published complete proteomes often have many sequences that are partial or derive from pseudogenes. Not only do these sequences add noise to the MSA, phylogeny and other downstream computational analyses, they also instigate many errors in the processing of the MSAs and downstream analyses, including the phylogeny. ObjectiveThis work aims to provide and test an objective, automated but flexible pipeline for the scrutiny of sequence sets from large, complex, eukaryotic protein superfamilies. The pipeline should classify sequences with high precision and recall as either functional or non-functional. The pipeline should classify no or only a few SwissProt sequences as non-functional (high precision) and sequences from other related superfamilies as non-functional (high recall) and result in a demonstrably much improved MSA (high performance). ResultsSeqrutinator is a pipeline that consists of five modules written in Python3 that identify and remove sequences that are likely Non-Functional Homologues (NFH). Here we tested the pipeline using three complex plant superfamilies (BAHD, CYP and UGT) that act in specialized metabolism, using the complete proteomes of 16 plant species as input and SwissProt as a control. Only 1.94% of SwissProt sequences with wetlab evidence were identified as NFH and all sequences from other related superfamilies were removed. Most NFH sequences are partial but, interestingly, their removal results in highly improved MSAs. a few but significant sequences that instigate large gaps were found. The five modules show similar behaviour when applied to the 16 sequence sets of the three analysed superfamilies. Pipelines with different module orders result in similar classifications and, moreover, show that different modules often detect the same sequences. Conclusion and perspectiveSeqrutinator forms a consistent pipeline for sequence scrutiny that does result in sequence sets that generate high fidelity MSAs. Recovery analyses show the method has high precision and recall.

bioinformatics↗