bioRxiv Science⌕ Search

Biology subjects

Roman, E. A.

Publications and source records attributed to Roman, E. A..

2 recordsLinked to original sources

PhISCO: a simple method to infer phenotypes from protein sequences

Although protein sequences encode the information for folding and function, understanding their link is not an easy task. Unluckily, the prediction of how specific amino acids contribute to these features is still considerably impaired. Here, we developed PhISCO, Phenotype Inference from Sequence COmparisons, a simple algorithm that finds positions associated with any quantitative phenotype and predicts their values. From a few hundred sequences from four different protein families, we performed multiple sequence alignments and calculated per-position pairwise differences for both the sequence and the observed phenotypes. We found that from 3 to 10 positions, depending on the studied case, were enough to identify positions associated with the phenotypes and perform quantitative predictions of them. Here we show that these strong correlations can be found using individual positions while an improvement is achieved when the most correlated positions are jointly analyzed. Noteworthy, we performed phenotype predictions using a simple linear model that links per-position divergences and differences in observed phenotypes. We also show that although extremely simple, predictions are comparable to the state-of-art methodologies which, in most of the cases, are far more complex. All of the calculations are obtained at a very low information cost since the only input needed is a multiple sequence alignment of protein sequences with their associated quantitative phenotype. The diversity of the explored systems makes PhISCO a valuable tool to find sequence determinants of biological activity modulation and to predict various functional features for uncharacterized members of a protein family.

biophysics↗

The N-terminal domain of RfaH plays an active role in protein fold-switching

The bacterial elongation factor RfaH promotes the expression of virulence factors by specifically binding to RNA polymerases (RNAP) stalled at a DNA signal known as ops. This behavior is unlike that of its paralog NusG, the major representative of the protein family to which RfaH belongs. Both proteins have an N-terminal domain (NTD) bearing an RNAP binding site, yet NusG C-terminal domain (CTD) is folded as a {beta}-barrel while RfaH CTD is forming an -hairpin blocking such site. Upon recognition of the ops exposed by RNAP, RfaH is activated via interdomain dissociation and complete CTD structural rearrangement into a {beta}-barrel structurally identical to NusG CTD. Although RfaH transformation has been extensively characterized computationally, most studies employ tertiary biases towards each native state, hampering the analysis of sequence-encoded interactions on fold-switching. Here, we used Associative Water-mediated Structure and Energy Model (AWSEM) molecular dynamics to characterize the transformation of RfaH, spotlighting the sequence-dependent effects of NTD on CTD fold stabilization. Umbrella sampling simulations guided by native contacts recapitulate the thermodynamic equilibrium experimentally observed for RfaH and its isolated CTD. Temperature refolding simulations of full-length RfaH show a high success towards -folded CTD, whereas the NTD interferes with {beta}CTD folding, becoming trapped in a {beta}-barrel intermediate. Meanwhile, NusG CTD refolding is unaffected by the presence of RfaH NTD, showing that these NTD-CTD interactions are encoded in RfaH sequence. Altogether, these results suggest that the NTD of RfaH favors the -folded RfaH by specifically orienting the CTD upon interdomain binding and also by favoring {beta}-barrel rupture into an intermediate from which fold-switching proceeds.

biophysics↗