bioRxiv Science⌕ Search

Biology subjects

Chauveau, M.

Publications and source records attributed to Chauveau, M..

2 recordsLinked to original sources

Improved inference of multiscale sequence statistics in generative protein models

High dimensionality and multiscale statistical structure are pervasive features of biological data, posing fundamental challenges for modeling. Because model inference generally proceeds with far fewer data than parameters, statistical patterns across scales are often unevenly represented. Protein sequences provide a paradigmatic example: statistics across homologs are inherently multiscale, displaying collective correlations among conserved residue sectors that encode function, alongside localized correlations corresponding to physical contacts outside these sectors. Standard regularization strategies used to mitigate undersampling during model inference have been shown to capture these patterns unevenly, a bias that compromises generative models of protein sequences by limiting their ability to produce both functional and diverse proteins. This limitation is exemplified by Boltzmann Machine-based generative models, which so far have required post hoc corrections to recover functionality, at the cost of reduced sequence diversity and novelty. Here, we introduce the stochastic Boltzmann Machine (sBM), a new regularization strategy that more accurately captures different correlation scales. Through analyses of theoretical models with known ground-truth parameters and experiments on the chorismate mutase family, we show that sBM effectively mitigates distortions in the estimation of model parameters, enabling the generation of functional sequences with greater diversity and without the need for post hoc corrections. These results advance the inference of generative models that more faithfully reflect the evolutionary constraints shaping protein sequences.

systems biology↗

COCOA-Tree: Phylogenetic visualization and comparative analysis of coevolving residues

The evolutionary co-occurrence of amino acid changes between protein residues underlies key structural and functional properties of protein families. Building on these coevolutionary patterns, methods have been developed to identify groups of residues associated with enzyme functionalities, such as Statistical Coupling Analysis (SCA) or Specificity-Determining Position (SDP) methods. These methods and their variations differ in the metrics used to quantify coevolution, residues weighting schemes, and corrections introduced to mitigate noise and phylogenetic biases. Yet, systematic comparisons across methods are rarely performed, and the evolutionary origins of the coevolutionary patterns highlighted by each approach are seldom addressed, limiting our ability to disentangle functional from phylogenetic contributions. To address these issues, we introduce COCOA-Tree, a Python library for SCA-like dimensionality-reduction analyses. COCOA-Tree supports custom metrics and enables visualization of coevolutionary patterns on phylogenetic trees. We also provide guidance to map results onto 3D structures in PyMOL. Using COCOA-Tree, we reanalyze published datasets and uncover previously unnoticed evolutionary properties of groups of coevolving residues detected by SCA, known as sectors. In particular, in the well-studied S1A serine protease family, we show that two of the three known sectors exhibit qualitatively distinct levels of sequence conservation depending on the enzymatic functions and on the phylogenetic clades to which the proteins belong. We further show that different coevolution metrics often identify qualitatively distinct groups of coevolving residues, although they yield consistent results for mildly conserved residues. Finally, as an example of the versatility of COCOA-Tree, we provide an example of visualizing pairs of residues detected by Direct Coupling Analysis (DCA) methods on a phylogenetic tree, highlighting a rich diversity of co-evolutionary patterns. Overall, we expect COCOA-Tree to help identify residues that control protein function and thereby improve our capacity for functional engineering and our understanding of the principles governing protein evolution. COCOA-Tree website: https://tree-timc.github.io/cocoatree

bioinformatics↗