bioRxiv Science⌕ Search

Biology subjects

Faldynova, H.

Publications and source records attributed to Faldynova, H..

2 recordsLinked to original sources

Uncovering Functional Distant Mutations by Ultra-High-Throughput Screening of Dehalogenases

Conformational dynamics play a central role in enzyme function by controlling substrate access and productive binding. Yet mutations that beneficially modulate these properties are difficult to identify. Here, we used ultrahigh-throughput fluorescence-activated droplet sorting (FADS) with a bulky fluorogenic substrate derived from coumarin (COU-3) to impose steric selection pressure on the haloalkane dehalogenase LinB. Screening a focused library yielded five single substitutions located 11.5-15.5 [A] from the catalytic centre. Variant I138N showed a fourfold increase in catalytic efficiency toward COU-3 through reduced KM and increased kcat, associated with increased cap-domain flexibility and facilitated substrate entry. In contrast, variant P208S markedly reduced substrate inhibition and shifted specificity toward bulkier iodinated haloalkanes by reshaping its tunnel environment. Integrated kinetic and structural analyses revealed that screening with bulky substrates directs selection toward distal regions controlling substrate access and unproductive binding. These findings demonstrate that ultrahigh-throughput FADS can reveal dynamic mechanisms of enzyme adaptation that remain difficult to predict by rational design. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=183 SRC="FIGDIR/small/713925v1_ufig1.gif" ALT="Figure 1"> View larger version (51K): org.highwire.dtl.DTLVardef@782038org.highwire.dtl.DTLVardef@8b43f3org.highwire.dtl.DTLVardef@11a403eorg.highwire.dtl.DTLVardef@6fcaea_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry↗

SoluProtMut: Siamese Deep Learning for Predicting Solubility Effects of Protein Mutations with Experimental Validation

Protein solubility is an attractive engineering target because it is a critical property influencing the scalability of protein production and the success of therapeutic proteins in biomedical applications. However, predicting solubility changes upon mutation in silico is challenging due to data heterogeneity and protein bias. Here, we explore how different sources of solubility data can be used for machine learning and present SoluProtMut, a Siamese deep geometric neural network trained to predict the impact of mutations on protein solubility. Our final model was trained exclusively on deep mutational scanning data. We compare our model with five established solubility prediction methods. The model achieves state-of-the-art performance on an independent dataset of various proteins, especially in predicting the effects of multipoint mutations (informedness of 26.5 %). Our findings also reaffirm that the scarcity of solubility data continues to hamper progress in this field. To address this limitation, we experimentally quantified solubility changes for hundreds of single-point and multipoint mutants of haloalkane dehalogenase. Complemented with recent deep-mutational-scanning data on myoglobin, we employed both these data for external validation. Although the generalization to unseen proteins remains limited, our findings demonstrate the potential of integrating high-throughput assays with deep learning to improve the accuracy and scope of solubility prediction. Highlights- We present a novel anti-symmetric Siamese architecture for mutational prediction on structures built on graph-based convolutional neural networks ensuring SE(3) invariance. - We demonstrate that a model trained only on single-point mutants of a single protein derived from high-throughput experiments generalizes to multipoint mutants of unseen proteins, achieving the state-of-the-art binary prediction informedness of 26.5 %. - We address the key limitation in the domain by extending the available data with 277 single-point and multipoint mutants of haloalkane dehalogenase labelled in-house and a selection of recently published 1037 single-point mutants of myoglobin. - We systematically study how different data subsets affect the performance of the trained models, revealing that including yeast-derived high-throughput data in training hampers generalization to low-throughput assays but recovers the performance on the yeast-derived myoglobin dataset.

bioinformatics↗