bioRxiv · 10.1101/2024.12.04.626803
Linguistic networks uncover grammatical constraints of protein sentences comprised by domain-based words
Abstract
Protein-protein interaction networks can help identify co-regulated modules, emergent biology (such as pathways), disease-associated partners, and function through association. However, these networks are limited by the breadth of experimental data behind them, which is incomplete and uneven across the proteome. For example, the coverage of interactions driven by reversible post-translational modifications (acetylation, phosphorylation, etc.) or protein fusions in diseases are particularly difficult to establish. Protein domains, conserved structural and functional units, are a major component of what defines a proteins function and its interactions. Whereas protein interactions are sparsely understood at this time, domain identification and coverage is mature. In this work, we propose a language-based network that utilizes the domain as "words" that makes up protein "sentences" to cover an entire proteome. We first convert proteins into n-grams (a formalization of contiguous words in a sentence) and assemble n-grams into a comprehensive network. We then use information theory to reduce the complexity of the network, collapsing to a network and n-gram size that recovers the majority of the proteome complexity. Using network theoretic approaches, we explore the larger human proteome and subnetwork analysis to understand specific properties in two applications: reversible systems of post-translational modifications (PTMs) and cancer fusion genes. PTM analysis across species suggests that reversible PTM systems convergently evolved similar domain architectures - allowing higher interconnectivity between reader and writer domains, while eraser domains remained highly disconnected. Cancer fusion analysis finds that, despite the possibility that fusions may sample novel domain word combinations, creating new connections or altering the human protein network, most cancer fusion genes follow existing domain combination rules. Collectively, these results suggest that an n-gram based analysis of proteomes complements direct protein interaction approaches, but provides a more fully described network of interconnected protein function that can provide unique insights on signaling pathway analysis. We refer to this approach for converting proteomes into functional linguistic networks as Domain Architecture Network Syntax (DANSy). SignificanceHere, we develop a novel computational method that treats proteins as sentences made up of domain words. We integrate this linguistic representation with networks to uncover abstract functional relationships to more fully represent the proteome than traditional protein-protein interaction networks, which are limited by incomplete experimental data. Our findings show how this framework, which we term Domain Architecture Network Syntax (DANSy), uncovers common "grammar" governing protein functionality in the evolution of post-translational modification systems and demonstrates how cancer fusion genes maintain established principles of the natural proteome.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shimpi, A. A., Naegle, K. M.. 2024-12-04. Linguistic networks uncover grammatical constraints of protein sentences comprised by domain-based words. https://doi.org/10.1101/2024.12.04.626803
Cite the original work for its findings. Save a collection to share your selection of sources.