bioRxiv Science⌕ Search

bioRxiv · 10.1101/2022.12.12.520004

Infer global, predict local: quantity-quality trade-off in protein fitness predictions from sequence data

Abstract

Predicting the effects of mutations on protein function is an important issue in evolutionary biology and biomedical applications. Computational approaches, ranging from graphical models to deep-learning architectures, can capture the statistical properties of sequence data and predict the outcome of high-throughput mutagenesis experiments probing the fitness landscape around some wild-type protein. However, how the complexity of the models and the characteristics of the data combine to determine the predictive performance remains unclear. Here, based on a theoretical analysis of the prediction error, we propose descriptors of the sequence data, characterizing their quantity and quality relative to the model. Our theoretical framework identifies a trade-off between these two quantities, and determines the optimal subset of data for the prediction task, showing that simple models can outperform complex ones when inferred from adequately-selected sequences. We also show how repeated subsampling of the sequence data allows for assessing how much epistasis in the fitness landscape is not captured by the computational model. Our approach is illustrated on several protein families, as well as on in silico solvable protein models. Significance StatementIs more data always better? Or should one prefer fewer data, but of higher quality? Here, we investigate this question in the context of the prediction of fitness effects resulting from mutations to a wild-type protein. We show, based on theory and data analysis, that simple models trained on a small subset of carefully chosen sequence data can perform better than complex ones trained on all available data. Furthermore, we explain how comparing the simple local models obtained with different subsets of training data reveals how much of the epistatic interactions shaping the fitness landscape are left unmodeled.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Posani, L., Rizzato, F., Monasson, R., Cocco, S.. 2022-12-14. Infer global, predict local: quantity-quality trade-off in protein fitness predictions from sequence data. https://doi.org/10.1101/2022.12.12.520004

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Scaling of structural variability of ecDNA polymer condensates with copy number boosts and stabilises oncogene regulatory contacts

Extrachromosomal DNAs (ecDNAs) form highly heterogeneous condensates in cancer cells that drive oncogene overexpression, yet how structural variability coexists with stable gene regulation remains unclear. Here, we develop a minimal polymer physics model of MYC-harbouring COLO320-DM ecDNAs, where BRD4-like complexes bind and bridge cognate sites along ecDNA rings. Above a critical binder concentration, ecDNAs phase separate into condensates exhibiting diverse conformations because of their thermodynamic folding degeneracy. Despite this variability, condensates retain conserved interaction scaffolds that give rise to reproducible contact patterns, including in-trans associated domains (I-TADs), genomic regions enriched in intermolecular regulatory contacts between distinct ecDNAs. We find that condensate 3D architecture follows universal scaling relations with ecDNA copy number, n, remaining robust to model parameter changes. Regulatory contacts within I TADs increase linearly with n, yet they are one order of magnitude stronger than in size matched control regions outside I TADs, whereas their relative fluctuations are markedly suppressed as n increases. This scaling produces enhanced, low-noise regulatory environments for oncogenes embedded within I-TADs, such as PVT1-MYC fusions, whereas the canonical MYC copy, located outside, is less amplified as experimentally observed. Our findings reveal universal polymer physics principles underlying ecDNA condensate organization, offering a mechanistic basis for selective oncogene amplification and potential advantages in cancer progression.

biophysics↗

High-resolution mapping of RNA structural maturation during Cas9 assembly with ABEL-FRET

The structural flexibility of RNA is essential for forming ribonucleoprotein (RNP) complexes, which regulate diverse biological processes. This intrinsic property permits RNA to act as a dynamic scaffold along the assembly pathway as it folds into a specific structure for initial recognition by protein and undergoes conformational rearrangements for functional maturation as a complex. Yet, RNA flexibility and RNP multicomponent assembly create significant obstacles for traditional structural methods. To overcome these challenges, we applied recently developed ABEL-FRET spectroscopy to measure tether-free single-molecule Forster resonance energy transfer (smFRET) over extended observation times. Furthermore, ABEL-FRET enables the unique ability for simultaneous measurements of ultrahigh resolution smFRET and hydrodynamic size of individual complexes, which offers distinct advantages for studying dynamic RNA molecules that undergo assembly via sequential binding events. Using ABEL-FRET, we explored how the guide RNA (gRNA) of CRISPR genome editing system folds and modulates its structural flexibility to carry out the roles required for each assembly state from its unbound apo form to the functional Cas9 RNP state for target DNA cleavage. Multi-perspective view of gRNA structure gained by probing its two primary functional domains enabled to capture dramatic changes in gRNA flexibility that are highly dependent on its specific structural domains as well as assembly states. Collectively, our work with ABEL-FRET highlights the intrinsic link between the structural flexibility of RNA and its functionality in RNP assembly.

biophysics↗

De novo design of functional RNAs through higher-order interactions

Designing RNA sequences that reliably adopt functional three-dimensional structures remains a central challenge in RNA engineering because folding depends on cooperative interactions beyond canonical base pairing. Here we present DS3dRNA, an interaction-based framework for de novo RNA sequence design that combines a three-body statistical potential with physics-guided sequence sampling and supports design against multiple conformations. Across the evaluated benchmarks, DS3dRNA outperformed representative RNA inverse-design methods in native-sequence recovery and agreement between predicted and target structures. Energy-sequence-quality analyses further showed that lower design energies generally accompanied higher sequence recovery and macro-averaged F1 scores (MacroF1). Experimentally tested Mango II designs retained high-affinity fluorogenic activity, and five twister ribozyme designs yielded mean endpoint cleavage fractions of 37.7-50.6%, compared with 23.5% for the wild type. These results establish explicit higher-order interaction scoring as a complementary approach to emerging data-driven RNA design methods and provide a framework for designing functional RNAs from experimental or predicted structural ensembles.

biophysics↗