bioRxiv Science⌕ Search

Biology subjects

Nisonoff, H.

Publications and source records attributed to Nisonoff, H..

4 recordsLinked to original sources

Eukaryotic RNA-guided endonucleases evolved from a unique clade of bacterial enzymes

RNA-guided endonucleases form the crux of diverse biological processes and technologies, including adaptive immunity, transposition, and genome editing. Some of these enzymes are components of insertion sequences (IS) in the IS200/IS605 and IS607 transposon families. Both IS families encode a TnpA transposase and TnpB nuclease, an RNA-guided enzyme ancestral to CRISPR-Cas12. In eukaryotes and their viruses, TnpB homologs occur as two distinct types, Fanzor1 and Fanzor2. We analyzed the evolutionary relationships between prokaryotic TnpBs and eukaryotic Fanzors, revealing that a clade of IS607 TnpBs with unusual active site arrangement found primarily in Cyanobacteriota likely gave rise to both types of Fanzors. The wide-spread nature of Fanzors imply that the properties of this particular group of IS607 TnpBs were particularly suited to adaptation and evolution in eukaryotes and their viruses. Experimental characterization of a prokaryotic IS607 TnpB and virally encoded Fanzor1s uncovered features that may have fostered coevolution between TnpBs/Fanzors and their cognate transposases. Our results provide insight into the evolutionary origins of a ubiquitous family of RNA-guided proteins that shows remarkable conservation across domains of life.

molecular biology↗

Discovery and validation of the binding poses of allosteric fragment hits to PTP1b: From molecular dynamics simulations to X-ray crystallography

Fragment-based drug discovery has led to six approved drugs, but the small size of the chemical fragments used in such methods typically results in only weak interactions between the fragment and its target molecule, which makes it challenging to experimentally determine the three-dimensional poses fragments assume in the bound state. One computational approach that could help address this difficulty is long-timescale molecular dynamics (MD) simulation, which has been used in retrospective studies to recover experimentally known binding poses of fragments. Here, we present the results of long-timescale MD simulations that we used to prospectively discover binding poses for two series of fragments in allosteric pockets on a difficult and important pharmaceutical target, protein-tyrosine phosphatase 1b (PTP1b). Our simulations reversibly sampled the fragment association and dissociation process. One of the binding pockets found in the simulations has not to our knowledge been previously observed with a bound fragment, and the other pocket adopted a very rare conformation. We subsequently obtained high-resolution crystal structures of members of each fragment series bound to PTP1b, and the experimentally observed poses confirmed the simulation results. To the best of our knowledge, our findings provide the first demonstration that MD simulations can be used prospectively to determine fragment binding poses to previously unidentified pockets.

biophysics↗

Combining evolutionary and assay-labelled data for protein fitness prediction

Predictive modelling of protein properties has become increasingly important to the field of machine-learning guided protein engineering. In one of the two existing approaches, evolutionarily-related sequences to a query protein drive the modelling process, without any property measurements from the laboratory. In the other, a set of protein variants of interest are assayed, and then a supervised regression model is estimated with the assay-labelled data. Although a handful of recent methods have shown promise in combining the evolutionary and supervised approaches, this hybrid problem has not been examined in depth, leaving it unclear how practitioners should proceed, and how method developers should build on existing work. Herein, we present a systematic assessment of methods for protein fitness prediction when evolutionary and assay-labelled data are available. We find that a simple baseline approach we introduce is competitive with and often outperforms more sophisticated methods. Moreover, our simple baseline is plug-and-play with a wide variety of established methods, and does not add any substantial computational burden. Our analysis highlights the importance of systematic evaluations and sufficient baselines.

synthetic biology↗

Sparse Epistatic Regularization of Deep Neural Networks for Inferring Fitness Functions

Despite recent advances in high-throughput combinatorial mutagenesis assays, the number of labeled sequences available to predict molecular functions has remained small for the vastness of the sequence space combined with the ruggedness of many fitness functions. Expressive models in machine learning (ML), such as deep neural networks (DNNs), can model the nonlinearities in rugged fitness functions, which manifest as high-order epistatic interactions among the mutational sites. However, in the absence of an inductive bias, DNNs overfit to the small number of labeled sequences available for training. Herein, we exploit the recent biological evidence that epistatic interactions in many fitness functions are sparse; this knowledge can be used as an inductive bias to regularize DNNs. We have developed a method for sparse epistatic regularization of DNNs, called the epistatic net (EN), which constrains the number of non-zero coefficients in the spectral representation of DNNs. For larger sequences, where finding the spectral transform becomes computationally intractable, we have developed a scalable extension of EN, which subsamples the combinatorial sequence space uniformly inducing a sparse-graph-code structure, and regularizes DNNs using the resulting greedy optimization method. Results on several biological landscapes, from bacterial to protein fitness functions, show that EN consistently improves the prediction accuracy of DNNs and enables them to outperform competing models which assume other forms of inductive biases. EN estimates all the higher-order epistatic interactions of DNNs trained on massive sequence spaces--a computational problem that takes years to solve without leveraging the epistatic sparsity in the fitness functions.

bioinformatics↗