bioRxiv Science⌕ Search

Biology subjects

Wesp, V.

Publications and source records attributed to Wesp, V..

3 recordsLinked to original sources

Strong correlation between amino acid frequency and codon degeneracy in genetic codes across all domains of life

Since the discovery of the genetic code, a frequently discussed question is whether the numbers of synonymous codons for the various amino acids are randomly distributed or were shaped by evolutionary constraints. In this study, we analyze for the standard as well as alternative genetic codes, the correlations between codon degeneracy and amino acid frequencies in proteins (neglecting differences in gene expression). To taking into account the effect of GC content, expected codon frequencies rather than codon multiplicity need to be considered. A strong correlation of these frequencies with amino acid abundance is revealed. Furthermore, we identify consistent patterns of over- and underrepresentation of amino acids across domains of cellular life as well as viruses. For example, the codons for glutamate, aspartate, lysine, and methionine are consistently overrepresented across domains, while cysteine, arginine, histidine, and proline are underrepresented. Subsequently, we discuss the role of biosynthesis costs of amino acids and other factors such as the order of amino acid recruitment in early evolution, exposure to oxidative stress, and sulphur availability. We hypothesize that in a first phase of evolution, the genetic code evolved in a way so as to comply with the different demands for amino acids. In a second phase, after the code was frozen, changes in demand led to deviations from the strong correlation between codon multiplicity and amino acid frequency by natural selection. This comprehensive analysis offers new insights into the interplay between genetic code structure and amino acid usage across all domains of life.

genomics↗

Proteomic study for the prediction of μCT imaging with iodine

Iodine-based staining techniques are commonly used in histological imaging and micro-computed tomography ({micro}CT) due to iodines affinity for binding to specific molecules. However, the basis for tissue-specific contrast has not yet been sufficiently explored. In this study, we analyse the human proteome at four different levels: individual proteins, protein families, tissues with additional expression values for selected proteins, and organs as a distinct combination of different tissues. At each level, we try to identify proteins/groups with high potential for iodine binding, especially those rich in aromatic heterocyclic amino acids. Using bioinformatic methods, we evaluate the occurrence of aromatic/non-aromatic heterocyclic, carbocyclic, and the remaining 15 amino acids in 20,650 proteins, 1,487 families, 57 tissues, and 16 organs. At the protein level, structural proteins such as titin, nebulin, obscurin, mucin, filaggrin and hornerin have a high absolute number of aromatic heterocyclic amino acids, which could explain the high {micro}CT contrast in muscle, skin and mucosal tissues. At the next level, however, structural families (such as the Laminin-family) rank significantly lower in comparison. These results are reflected in tissues and organs for which protein expression is available. Here, no significant correlations between the enrichment of heterocycles and the intensity of iodine staining can be observed. Furthermore, the enrichment of amino acids in each tissue/organ is relatively similar and shows no significant difference. Our results provide a general basis for iodine-based tissue imaging and serve as a potential starting point for future research, e.g. for cross-species applications and for the structural and functional effects of iodination.

biochemistry↗

Evaluating interchain hydrogen bonds between collagens- A multiple alignment approach

Collagens are structural proteins that are predominantly found in the extracellular matrix of multicellular animals, where they are mainly responsible for the stability and structural integrity of various tissues. All collagens contain polypeptide strands ([a]-chains). There are several types of collagens, some of which differ significantly in form, function, and tissue specificity. Because of their importance in clinical research, they are grouped into subdivisions, the so-called collagen families, and their sequences are often analysed. However, problems arise with highly homologous sequence segments. To increase the accuracy of collagen classification and prediction of their functions, the structure of these collagens and their expression in different tissues could result in a better focus on sequence segments of interest. Here, we analyse collagen families with different levels of conservation. As a result, clusters with high interconnectivity can be found, such as the fibrillar collagens, the COL4 network-forming collagens, and the COL9 FACITs. Furthermore, a large cluster between network-forming, FACIT, and COL28a1 [a]-chains is formed with COL6a3 as a major hub node. The formation of clusters also signifies, why it is important to always analyse the [a]-chains and why structural changes can have a wide range of effects on the body.

systems biology↗