bioRxiv Science⌕ Search

Biology subjects

Rousset, Y.

Publications and source records attributed to Rousset, Y..

2 recordsLinked to original sources

Predicting turnover number fold-changes to recover true mutation effects and overcome biases in mutant datasets

Machine learning is increasingly used to guide protein engineering by predicting how mutations affect desired properties. Recent models for the turnover number (kcat) of enzymes report high accuracy, suggesting that mutation effects can be inferred directly from protein sequence. However, these approaches are typically evaluated on heterogeneous datasets of enzyme variants, where closely related sequences and systematic reporting patterns may confound model performance. A central challenge is therefore to determine whether current models truly capture mutation-specific effects or instead exploit statistical regularities in the data. Here we show that much of the reported accuracy in mutant kcat prediction arises from two pervasive biases: variants of the same enzyme occupy a narrow activity range, and mutations within a group often share a common direction of change. Simple baselines that exploit these biases match or exceed the performance of existing models, indicating that high apparent accuracy does not imply mechanistic understanding. To address this limitation, we introduce a bias-aware framework that reformulates prediction as a pairwise fold-change task and evaluates performance on unseen mutant-mutant pairs, thereby isolating mutation-specific signal. A proof-of-principle implementation explains approximately one-third of the variance under these conditions and outperforms existing models on leakage-controlled benchmarks. More broadly, this work establishes a general strategy for evaluating and modeling mutation effects in biochemical datasets, with implications for protein engineering and related fields.

bioinformatics↗

Stochastic modelling of a three-dimensional glycogen granule synthesis and impact of the branching enzyme

In humans, glycogen storage diseases result from metabolic inborn errors, and can lead to severe phenotypes and lethal conditions. Besides these rare diseases, glycogen is also associated to widely spread societal burdens such as diabetes. Glycogen is a branched glucose polymer synthesised and degraded by a complex set of enzymes. Over the past 50 years, the structure of glycogen has been intensively investigated. Yet, the interplay between glycogen structure and the related enzymes is still to be characterised. In this article, we develop a stochastic coarse-grained and spatially resolved model of branched polymer biosynthesis following a Gillespie algorithm. Our study largely focusses on the role of the branching enzyme, and first investigates the properties of the model with generic parameters, before comparing it to in vivo experimental data in mice. It arises that the ratio of glycogen synthase over branching enzyme activities drastically impacts the structure of the granule. We deeply investigate the mechanism of branching and parametrise it using distinct lengths. Not only do we consider various possible sets of values for these lengths, but also distinct rules to apply them. We show how combining them finely tunes glycogen macromolecular structure. Comparing the model with experimental data confirms that we can accurately reproduce glycogen chain length distributions in wild type mice. Additional granule properties obtained for this fit are also in good agreement with typically reported values in the experimental literature. Nonetheless, we find that the mechanism of branching must be more flexible than usually reported. Overall, we demonstrate that the chain length distribution is an imprint of the branching activity and mechanism. Our generic model and methods can be applied to any glycogen data set, and could in particular contribute to characterise the mechanisms responsible for glycogen storage disorders. Author summaryGlycogen is a granule-like macromolecule made of 10,000 to 50,000 glucose units arranged in linear and branched chains. It serves as energy storage in many species, including humans. Depending on physiological conditions (hormone concentrations, glucose level, etc.) glycogen granules are either synthesised or degraded. Certain metabolic disorders are associated to abnormal glycogen structures, and structural properties of glycogen might impact the dynamics of glucose release and storage. To capture the complex interplay between this dynamics and glycogen structural properties, we propose a computational model relying on the random nature of biochemical reactions. The granule is represented in three dimensions and resolved at the glucose scale. Granules are produced under the action of a complex set of enzymes, and we mostly focus on those responsible for the formation of new branches. Specifically, we study the impact of their molecular action on the granule structure. With this model, we are able to reproduce structural properties observed under certain in-vivo conditions. Our biophysical and computational approach complements experimental studies and may contribute to characterise processes responsible for glycogen related disorders.

biophysics↗