bioRxiv Science⌕ Search

Biology subjects

Spranger, L.

Publications and source records attributed to Spranger, L..

2 recordsLinked to original sources

The Role of Metabolism in Shaping Enzyme Structures Over 400 Million Years of Evolution

The functions of cells and proteins depend on their biochemical microenvironment. To understand how biochemical constraints shaped protein structural evolution, we coupled the extensive genetic and metabolic data from the Saccharomycotina subphylum with the capability of AlphaFold2 to systematically predict protein structures from sequence. Determining how 11,269 enzyme structures catalysing 361 different metabolic reactions evolved over 400 million years alongside their molecular functions, we report that metabolism has shaped the structural evolution of enzymes at different levels: the organisms overall metabolism; the topological organisation of the metabolic network; and each enzymes molecular properties. For example, structural evolution depends on each enzymes reaction mechanism, on the variability rather than the amount of metabolic flux, and on biosynthetic cost. Evolutionary cost-optimization is stronger on highly abundant enzymes and acts differently on different structural domains, with the exception of small-molecule binding sites, which are prioritised over other structural domains and lack cost-optimisation. Finally, while enzyme surfaces are less constrained, surface residues can also be exposed to positive selection for the co-evolution of protein-protein interaction sites. Accessing AlphaFolds power to predict protein structures systematically and across species barriers, facilitating the integration of protein structures with functional genomics, we were thus able to map biological constraints which shape protein structural evolution at scale and over long timelines.

systems biology↗

Node-degree aware edge sampling mitigates inflated classification performance in biomedical graph representation learning

Graph representation learning is a family of related approaches that learn low-dimensional vector representations of nodes and other graph elements called embeddings. Embeddings approximate characteristics of the graph and can be used for a variety of machine-learning tasks such as novel edge prediction. For many biomedical applications, partial knowledge exists about positive edges that represent relationships between pairs of entities, but little to no knowledge is available about negative edges that represent the explicit lack of a relationship between two nodes. For this reason, classification procedures are forced to assume that the vast majority of unlabeled edges are negative. Existing approaches to sampling negative edges for training and evaluating classifiers do so by uniformly sampling pairs of nodes. We show here that this sampling strategy typically leads to sets of positive and negative edges with imbalanced edge degree distributions. Using representative homogeneous and heterogeneous biomedical knowledge graphs, we show that this strategy artificially inflates measured classification performance. We present a degree-aware node sampling approach for sampling negative edge examples that mitigates this effect and is simple to implement.

bioinformatics↗