bioRxiv Science⌕ Search

Biology subjects

Aptekmann, A. A.

Publications and source records attributed to Aptekmann, A. A..

4 recordsLinked to original sources

NAC1 directs CEP1-CEP3 peptidase expression and decreases cell wall extensins linked to root hair growth in Arabidopsis

Plant genomes encode a unique group of papain-type Cysteine EndoPeptidases (CysEPs) containing a KDEL endoplasmic reticulum (ER) retention signal (KDEL-CysEPs or CEPs). CEPs process the cell-wall scaffolding EXTENSIN proteins (EXTs), which regulate de novo cell wall formation and cell expansion. Since CEPs are able to cleave EXTs and EXT-related proteins, acting as cell wall-weakening agents, they may play a role in cell elongation. Arabidopsis thaliana genome encodes three CEPs (AtCPE1-AtCEP3). Here we report that the three Arabidopsis CEPs, AtCEP1-AtCEP3, are highly expressed in root-hair cell files. Single mutants have no evident abnormal root-hair phenotype, but atcep1-3 atcep3-2 and atcep1-3 atcep2-2 double mutants have longer root hairs (RHs) than wild type (Wt) plants, suggesting that expression of AtCEPs in root trichoblasts restrains polar elongation of the RH. We provide evidence that the transcription factor NAC1 activates AtCEPs expression in roots to limit RH growth. Chromatin immunoprecipitation indicates that NAC1 binds the promoter of AtCEP1, AtCEP2, and to a lower extent to AtCEP3 and may directly regulate their expression. Indeed, inducible NAC1 overexpression increases AtCEP1 and AtCEP2 transcript levels in roots and leads to reduced RH growth while the loss of function nac1-2 mutation reduces AtCEP1-AtCEP3 gene expression and enhances RH growth. Likewise, expression of a dominant chimeric NAC1-SRDX repressor construct leads to increased RH length. Finally, we show that RH cell walls in the atcep1-1 atcep3-2 double mutant have reduced levels of EXT deposition, suggesting that the defects in RH elongation are linked to alterations in EXT processing and accumulation. Taken together, our results support the involvement of AtCEPs in controlling RH polar growth through EXT-processing and insolubilization at the cell wall.

plant biology↗

PhISCO: a simple method to infer phenotypes from protein sequences

Although protein sequences encode the information for folding and function, understanding their link is not an easy task. Unluckily, the prediction of how specific amino acids contribute to these features is still considerably impaired. Here, we developed PhISCO, Phenotype Inference from Sequence COmparisons, a simple algorithm that finds positions associated with any quantitative phenotype and predicts their values. From a few hundred sequences from four different protein families, we performed multiple sequence alignments and calculated per-position pairwise differences for both the sequence and the observed phenotypes. We found that from 3 to 10 positions, depending on the studied case, were enough to identify positions associated with the phenotypes and perform quantitative predictions of them. Here we show that these strong correlations can be found using individual positions while an improvement is achieved when the most correlated positions are jointly analyzed. Noteworthy, we performed phenotype predictions using a simple linear model that links per-position divergences and differences in observed phenotypes. We also show that although extremely simple, predictions are comparable to the state-of-art methodologies which, in most of the cases, are far more complex. All of the calculations are obtained at a very low information cost since the only input needed is a multiple sequence alignment of protein sequences with their associated quantitative phenotype. The diversity of the explored systems makes PhISCO a valuable tool to find sequence determinants of biological activity modulation and to predict various functional features for uncharacterized members of a protein family.

biophysics↗

mebipred: identifying metal-binding potential in protein sequences.

Metal-binding proteins have a central role in maintaining life processes. Nearly one-third of known protein structures contain metal ions that are used for a variety of needs, such as catalysis, DNA/RNA binding, protein structure stability, etc. Identifying metal-binding proteins is thus crucial for understanding the mechanisms of cellular activity. However, experimental annotation of protein metal-binding potential is severely lacking, while computational techniques are often imprecise and of limited applicability. We developed a novel machine learning-based method, mebipred, for identifying metal-binding proteins from sequence-derived features. This method is nearly 90% accurate in recognizing proteins that bind metal ions and ion containing ligands. Moreover, the identity of ten ubiquitously present metal ions and ion-containing ligands can be annotated. mebipred is reference-free, i.e. no sequence alignments are involved, and outperforms other prediction methods, both in speed and accuracy. mebipred can also identify protein metal-binding capabilities from short sequence stretches and, thus, may be useful for the annotation of metagenomic samples metal requirements inferred from translated sequencing reads. We performed an analysis of microbiome data and found that ocean, hot spring sediments and soil microbiomes use a more diverse set of metals than human host-related ones. For human-hosted microbiomes, physiological conditions explain the observed metal preferences. Similarly, subtle changes in ocean sample ion concentration affect the abundance of relevant metal-binding proteins. These results are highlight mebipreds utility in analyzing microbiome metal requirements. mebipred is available as a web server at services.bromberglab.org/mebipred and as a standalone package at https://pypi.org/project/mymetal/

bioinformatics↗

Decoding the effects of synonymous variants

Synonymous single nucleotide variants (sSNVs) are common in the human genome but are often overlooked. However, sSNVs can have significant biological impact and may lead to disease. Existing computational methods for evaluating the effect of sSNVs suffer from the lack of gold-standard training/evaluation data and exhibit over-reliance on sequence conservation signals. We developed synVep (synonymous Variant effect predictor), a machine learning-based method that overcomes both of these limitations. Our training data was a combination of variants reported by gnomAD (observed) and those unreported, but possible in the human genome (generated). We used positive-unlabeled learning to purify the generated variant set of any likely unobservable variants. We then trained two sequential extreme gradient boosting models to identify subsets of the remaining variants putatively enriched and depleted in effect. Our method attained 90% precision/recall on a previously unseen set of variants. Furthermore, although synVep does not explicitly use conservation, its scores correlated with evolutionary distances between orthologs in cross-species variation analysis. synVep was also able to differentiate pathogenic vs. benign variants, as well as splice-site disrupting variants (SDV) vs. non-SDVs. Thus, synVep provides an important improvement in annotation of sSNVs, allowing users to focus on variants that most likely harbor effects.

bioinformatics↗