Strong correlation between amino acid frequency and codon degeneracy in genetic codes across all domains of life
Since the discovery of the genetic code, a frequently discussed question is whether the numbers of synonymous codons for the various amino acids are randomly distributed or were shaped by evolutionary constraints. In this study, we analyze for the standard as well as alternative genetic codes, the correlations between codon degeneracy and amino acid frequencies in proteins (neglecting differences in gene expression). To taking into account the effect of GC content, expected codon frequencies rather than codon multiplicity need to be considered. A strong correlation of these frequencies with amino acid abundance is revealed. Furthermore, we identify consistent patterns of over- and underrepresentation of amino acids across domains of cellular life as well as viruses. For example, the codons for glutamate, aspartate, lysine, and methionine are consistently overrepresented across domains, while cysteine, arginine, histidine, and proline are underrepresented. Subsequently, we discuss the role of biosynthesis costs of amino acids and other factors such as the order of amino acid recruitment in early evolution, exposure to oxidative stress, and sulphur availability. We hypothesize that in a first phase of evolution, the genetic code evolved in a way so as to comply with the different demands for amino acids. In a second phase, after the code was frozen, changes in demand led to deviations from the strong correlation between codon multiplicity and amino acid frequency by natural selection. This comprehensive analysis offers new insights into the interplay between genetic code structure and amino acid usage across all domains of life.