bioRxiv Science⌕ Search

Biology subjects

Boorla, V. S.

Publications and source records attributed to Boorla, V. S..

4 recordsLinked to original sources

CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters kcat, Km and Ki

Quantification of enzymatic activities still heavily relies on experimental assays, which can be expensive and time-consuming. Therefore, methods that enable accurate predictions of enzyme activity can serve as effective digital twins. A few recent studies have shown the possibility of training machine learning (ML) models for predicting the enzyme turnover numbers (kcat) and Michaelis constants (Km) using only features derived from enzyme sequences and substrate chemical topologies by training on in vitro measurements. However, several challenges remain such as lack of standardized training datasets, evaluation of predictive performance on out-of-distribution examples, and model uncertainty quantification. Here, we introduce CatPred, a comprehensive framework for ML prediction of in vitro enzyme kinetics. We explored different learning architectures and feature representations for enzymes including those utilizing pretrained protein language model features and pretrained three-dimensional structural features. We systematically evaluate the performance of trained models for predicting kcat, Km, and inhibition constants (Ki) of enzymatic reactions on held-out test sets with a special emphasis on out-of-distribution test samples (corresponding to enzyme sequences dissimilar from those encountered during training). CatPred assumes a probabilistic regression approach offering query-specific standard deviation and mean value predictions. Results on unseen data confirm that accuracy in enzyme parameter predictions made by CatPred positively correlate with lower predicted variances. Incorporating pre-trained language model features is found to be enabling for achieving robust performance on out-of-distribution samples. Test evaluations on both held-out and out-of-distribution test datasets confirm that CatPred performs at least competitively with existing methods while simultaneously offering robust uncertainty quantification. CatPred offers wider scope and larger data coverage ([~]23k, 41k, 12k data-points respectively for kcat, Km and Ki). A web-resource to use the trained models is made available at: https://tiny.cc/catpred

bioinformatics↗

A CNN model for predicting binding affinity changes between SARS-CoV-2 spike RBD variants and ACE2 homologues

The cellular entry of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) involves the association of its receptor binding domain (RBD) with human angiotensin converting enzyme 2 (hACE2) as the first crucial step. Efficient and reliable prediction of RBD-hACE2 binding affinity changes upon amino acid substitutions can be valuable for public health surveillance and monitoring potential spillover and adaptation into non-human species. Here, we introduce a convolutional neural network (CNN) model trained on protein sequence and structural features to predict experimental RBD-hACE2 binding affinities of 8,440 variants upon single and multiple amino acid substitutions in the RBD or ACE2. The model achieves a classification accuracy of 83.28% and a Pearson correlation coefficient of 0.85 between predicted and experimentally calculated binding affinities in five-fold cross-validation tests and predicts improved binding affinity for most circulating variants. We pro-actively used the CNN model to exhaustively screen for novel RBD variants with combinations of up to four single amino acid substitutions and suggested candidates with the highest improvements in RBD-ACE2 binding affinity for human and animal ACE2 receptors. We found that the binding affinity of RBD variants against animal ACE2s follows similar trends as those against human ACE2. White-tailed deer ACE2 binds to RBD almost as tightly as human ACE2 while cattle, pig, and chicken ACE2s bind weakly. The model allows testing whether adaptation of the virus for increased binding with other animals would cause concomitant increases in binding with hACE2 or decreased fitness due to adaptation to other hosts.

biophysics↗

Computational prediction of the effect of amino acid changes on the binding affinity between SARS-CoV-2 spike protein and the human ACE2 receptor

The association of the receptor binding domain (RBD) of SARS-CoV-2 viral spike with human angiotensin converting enzyme (hACE2) represents the first required step for viral entry. Amino acid changes in the RBD have been implicated with increased infectivity and potential for immune evasion. Reliably predicting the effect of amino acid changes in the ability of the RBD to interact more strongly with the hACE2 receptor can help assess the public health implications and the potential for spillover and adaptation into other animals. Here, we introduce a two-step framework that first relies on 48 independent 4-ns molecular dynamics (MD) trajectories of RBD-hACE2 variants to collect binding energy terms decomposed into Coulombic, covalent, van der Waals, lipophilic, generalized Born electrostatic solvation, hydrogen-bonding, {pi}-{pi} packing and self-contact correction terms. The second step implements a neural network to classify and quantitatively predict binding affinity using the decomposed energy terms as descriptors. The computational base achieves an accuracy of 82.2% in terms of correctly classifying single amino-acid substitution variants of the RBD as worsening or improving binding affinity for hACE2 and a correlation coefficient r of 0.69 between predicted and experimentally calculated binding affinities. Both metrics are calculated using a 5-fold cross validation test. Our method thus sets up a framework for effectively screening binding affinity change with unknown single and multiple amino-acid changes. This can be a very valuable tool to predict host adaptation and zoonotic spillover of current and future SARS-CoV-2 variants.

biophysics↗

De novo design of high-affinity antibody variable regions (scFv) against the SARS-CoV-2 spike protein

The emergence of SARS-CoV-2 is responsible for the pandemic of respiratory disease known as COVID-19, which emerged in the city of Wuhan, Hubei province, China in late 2019. Both vaccines and targeted therapeutics for treatment of this disease are currently lacking. Viral entry requires binding of the viral spike receptor binding domain (RBD) with the human angiotensin converting enzyme (hACE2). In an earlier paper1, we report on the specific residue interactions underpinning this event. Here we report on the de novo computational design of high affinity antibody variable regions through the recombination of VDJ genes targeting the most solvent-exposed hACE2-binding residues of the SARS-CoV-2 spike protein using the software tool OptMAVEn-2.02. Subsequently, we carry out computational affinity maturation of the designed prototype variable regions through point mutations for improved binding with the target epitope. Immunogenicity was restricted by preferring designs that match sequences from a 9-mer library of "human antibodies" based on H-score (human string content, HSC)3. We generated 106 different designs and report in detail on the top five that trade-off the greatest affinity for the spike RBD epitope (quantified using the Rosetta binding energies) with low H-scores. By grafting the designed Heavy (VH) and Light (VL) chain variable regions onto a human framework (Fc), high-affinity and potentially neutralizing full-length monoclonal antibodies (mAb) can be constructed. Having a potent antibody that can recognize the viral spike protein with high affinity would be enabling for both the design of sensitive SARS-CoV-2 detection devices and for their deployment as therapeutic antibodies.

immunology↗