bioRxiv · 10.64898/2025.12.31.697235
Proteins as Statistical Languages: Information-Theoretic Signatures of Proteomes Across the Tree of Life
Abstract
Protein sequences are commonly interpreted through biochemical and evolutionary lenses, emphasizing structure-function relationships and selection in sequence space. Here we develop a complementary viewpoint: proteins as statistical languages--strings over a finite alphabet generated by constrained stochastic processes. We formalize intrinsic informational descriptors of protein ensembles, including composition entropy H1, adjacent mutual information I1, and separation-dependent information profiles Id. A null-model ladder (uniform, composition-matched i.i.d., and Markov-1) separates compositional effects from genuine positional dependence. We then evaluate these descriptors empirically across 20 UniProt reference proteomes spanning major clades, using protein-level bootstrap resampling and matched synthetic controls. Real proteomes consistently depart from composition-matched i.i.d. baselines and exhibit information profiles that remain elevated beyond the decay expected under first-order Markov surrogates, indicating dependencies beyond local transition statistics. Finally, a compressibility proxy (gzip) provides an orthogonal signature of redundancy relative to i.i.d. controls at matched composition. Together, these results support the view of proteomes as constrained statistical languages and provide model-agnostic fingerprints for comparing sequence ensembles.These signatures provide a lightweight diagnostic layer for comparing proteomes prior to mechanistic modeling
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alegre, E. O. T.. 2026-01-05. Proteins as Statistical Languages: Information-Theoretic Signatures of Proteomes Across the Tree of Life. https://doi.org/10.64898/2025.12.31.697235
Cite the original work for its findings. Save a collection to share your selection of sources.