bioRxiv ScienceSearch

Biology subjects

Arne Elofsson

Publications and source records attributed to Arne Elofsson.

5 recordsLinked to original sources

Accurate contact predictions for thousands of protein families using PconsC3

Protein structure prediction was for decades one of the grand unsolved challenges in bioinformatics. A few years ago it was shown that by using a maximum entropy approach to describe couplings between columns in a multiple sequence alignment it was possible to significantly increase the accuracy of residue contact predictions. For very large protein families with more than 1000 effective sequences the accuracy is sufficient to produce accurate models of proteins as well as complexes. Today, for about half of all Pfam domain families no structure is known, but unfortunately most of these families have at most a few hundred members, i.e. are too small for existing contact prediction methods. To extend accurate contact predictions to the thousands of smaller protein families we present PconsC3, an improved method for protein contact predictions that can be used for families with as little as 100 effective sequence members. We estimate that PconsC3 provides accurate contact predictions for up to 4646 Pfam domain families. In addition, PconsC3 outperforms previous methods significantly independent on family size, secondary structure content, contact range, or the number of selected contacts. This improvement translates into improved de-novo prediction of three-dimensional structures. PconsC3 is available as a web server and downloadable version at http://c3.pcons.net. The downloadable version is free for all to use and licensed under the GNU General Public License, version 2.

Bioinformatics

High GC Content Causes De Novo Created Proteins to be Intrinsically Disordered

De novo creation of protein coding genes involves formation of short ORFs from noncoding regions; some of these ORFs might then become fixed in the population. De novo created proteins need to, at the bare minimum, not cause serious harm to the organism, meaning that they should for instance not cause aggregation. Therefore, although the creation of the short ORFs could be truly random, but the fixation should be of subject to some selective pressure. The selective forces acting on de novo created proteins have been elusive and contradictory results have been reported. In Drosophila they are more disordered, i.e. are enriched in polar residues, than ancient proteins, while the opposite trend is present in yeast. To the best of our knowledge no valid explanation for this difference has been proposed.\n\nTo solve this riddle we studied structural properties and age of all proteins in 187 eukaryotic species. We find that, on average, there are small differences between proteins of different ages, with the exception that younger proteins are shorter. However, when we take the GC content into account we find that this can explain the opposite trends observed in yeast (low GC) and drosophila (high GC). GC content is correlated with codons coding for disorder-promoting amino acids, and inversely correlated with transmembrane, helix and sheet promoting residues. We find that for the youngest proteins, i.e. the ones that are most likely to be de novo created, there exists a strong correlation with GC and structural properties. In contrast, this strong relationship is not seen for ancient proteins. This leads us to propose that structural features are not a strong determining factor for fixation of de novo created genes. Instead these proteins resemble random proteins given a particular GC level. The dependency on GC content is then gradually weakened during evolution.\n\nAuthor SummaryWe show that the GC content of a genomic area is of great importance for the properties of a protein-coding de novo created gene. The GC content affects the frequency of the codons and this affects the probability for each amino acid to be included in a de novo created protein. The codons encoding for Ala, Pro and Glu contain 80% GC, while codons for Lys, Phe, Asn, Tyr and Ile contain 20% or less. Pro and Gly are disorder-promoting, while Phe, Tyr and Ile are order-promoting. Therefore random protein sequences at a high GC will be more disordered than the ones created at a low GC. The structural properties of the youngest (orphan) proteins match to a large degree the properties of random proteins when the GC content is taken into account. In contrast structural properties of ancient proteins only show a weak correlation with GC content. This suggests that even after fixation of de novo created proteins largely resemble random proteins given a certain GC content. Thereafter, during evolution the correlation between structural properties and GC weakens.

Evolutionary Biology

Folding of Aquaporin 1: Multiple evidence that helix 3 can shift out of the membrane core

The folding of most integral membrane proteins follows a two-step process: Initially, individual transmembrane helices are inserted into the membrane by the Sec translocon. Thereafter, these helices fold to shape the final conformation of the protein. However, for some proteins, including Aquaporin 1 (AQP1), the folding appears to follow a more complicated path. AQP1 has been reported to first insert as a four-helical intermediate, where helix 2 and 4 are not inserted into the membrane. In a second step this intermediate is folded into a six-helical topology. During this process, the orientation of the third helix is inverted. Here, we propose a mechanism for how this reorientation could be initiated: First, helix 3 slides out from the membrane core resulting in that the preceding loop enters the membrane. The final conformation could then be formed as helix 2, 3 and 4 are inserted into the membrane and the reentrant regions come together. We find support for the first step in this process by showing that the loop preceding helix 3 can insert into the membrane. Further, hydrophobicity curves, experimentally measured insertion efficiencies and MD-simulations suggest that the barrier between these two hydrophobic regions is relatively low, supporting the idea that helix 3 can slide out of the membrane core, initiating the rearrangement process

Molecular Biology

The positive inside rule is stronger when followed by a transmembrane helix.

The translocon recognizes transmembrane helices with sufficient level of hydrophobicity and inserts them into the membrane. However, sometimes less hydrophobic helices are also recognized. Positive inside rule, orientational preferences of and specific interactions with neighboring helices have been shown to aid in the recognition of these helices, at least in artificial systems. To better understand how the translocon inserts marginally hydrophobic helices, we studied three naturally occurring marginally hydrophobic helices, which were previously shown to require the subsequent helix for efficient translocon recognition. We find no evidence for specific interactions when we scan all residues in the subsequent helices. Instead, we identify arginines located at the N-terminal part of the subsequent helices that are crucial for the recognition of the marginally hydrophobic transmembrane helices, indicating that the positive inside rule is important. However, in two of the constructs these arginines do not aid in the recognition without the rest of the subsequent helix, i.e. the positive inside rule alone is not sufficient. Instead, the improved recognition of marginally hydrophobic helices can here be explained as follows; the positive inside rule provides an orientational preference of the subsequent helix, which in turn allows the marginally hydrophobic helix to be inserted, i.e. the effect of the positive inside rule is stronger if positively charged residues are followed by a transmembrane helix. Such a mechanism can obviously not aid C-terminal helices and consequently we find that the terminal helices in multi-spanning membrane proteins are more hydrophobic than internal helices.

Molecular Biology

Accurate prediction of transmembrane β-barrel proteins from sequences

Transmembrane {beta}-barrels are known to play major roles in substrate transport and protein biogenesis in gram-negative bacteria, chloroplasts and mitochondria. However, the exact number of transmembrane {beta}-barrel families is unknown and experimental structure determination is challenging. In theory, if one knows the number of strands in the {beta}-barrel, then the 3D structure of the barrel could be trivial, but current topology predictions do not predict accurate structures and are unable to give information beyond the {beta}-strands in the barrel. Recent work has shown successful prediction of globular and alpha-helical membrane proteins from sequence alignments, by using high ranked evolutionary couplings between residues as distance constraints to fold extended polypeptides. However, these methods, have not addressed the calculation of precise {beta}-sheet hydrogen bonding that defines transmembrane {beta}-barrels, and would be required to fold these proteins successfully. Hence we developed a method (EVFold_BB) that can successfully model transmembrane {beta}-barrels by combining evolutionary couplings together with topology predictions. EVFold_BB is validated by the accurate all-atom 3D modeling of 18 proteins, representing all known membrane {beta}-barrel families that have sufficient sequences available. To demonstrate the potential of our approach we predict the unknown 3D structure of the LptD protein, the plausibility of its accuracy is supported by the blindly predicted benchmarks, and is consistent with experimental observations. Our approach can naturally be extended to all unknown {beta}-barrel proteins with sufficient sequence information.

Bioinformatics