bioRxiv Science⌕ Search

Biology subjects

Gaza, J.

Publications and source records attributed to Gaza, J..

3 recordsLinked to original sources

Efficient exploration of peptide libraries using active learning with AlphaFold-based screening

We previously showed that AlphaFold2 can be used to screen for peptide-binding epitopes targeting the extraterminal (ET) domain of Bromodomain and Extraterminal (BET) proteins from candidate protein partners identified in pull-down experiments. However, such approaches require large numbers of AlphaFold2 calculations, making exhaustive screening impractical for larger datasets, such as viral proteomes that may target the ET domain. In many cases, identifying a substantial fraction of binders--even without exhaustive coverage--would already provide valuable biological insight into these interaction networks. Here, we show that an active learning strategy based on Thompson sampling (TS) can efficiently explore peptide sequence space. Using a library derived from BRD3 pull-down experiments, TS recovers 50% of all binders using 15% of the queries required by exhaustive sampling (3.3 times improvement over random sampling). Moreover, TS consistently identifies experimentally known binding epitopes with substantially fewer queries. Because the approach relies only on binary labels, it is readily transferable to other protein-peptide systems where AF-based binding classification is applicable, as well as to peptide-property predictors for properties such as solubility or aggregation propensity.

bioinformatics↗

Hierarchical Extended Linkage Method (HELM)'s Deep Dive into Hybrid Clustering Strategies

Clustering remains a key tool in the analysis of molecular dynamics (MD) simulations, from the preparation of kinetic models to the study of mechanistic pathways and structural determination. It is no surprise then that multiple algorithms are currently used in the MD community, with k-means and hierarchical approaches being arguably the two most popular approaches. The former is very attractive from a purely computational point of view, demanding minimal memory and time resources, but at the price of being able to partition the data in very restrictive ways. Hierarchical strategies, on the other hand, can generate arbitrary partitions, but with steep memory and time requirements due to their need to build a pairwise distance matrix for all the considered conformations/frames. Here we propose a new hybrid paradigm, the Hierarchical Extended Linkage Method (HELM), that retains the efficiency of k-means while incorporating the flexibility of hierarchical methods. The key ingredient is the use of n-ary difference functions as a way to stabilize the k-means results and efficiently build the hierarchy of subsets. We showcase the applicability of this strategy over protein-DNA and protein folding studies, including the complete analysis of simulations with over 1.5 million frames. HELM is freely available in our MDANCE clustering package.

biophysics↗

PERRC: Protease Engineering with Reactant Residence Time Control

Proteases with engineered specificity hold great potential for targeted therapeutics, protein circuit construction, and biotechnology applications. However, many proteases exhibit broad substrate specificity, limiting their applications. Engineering protease specificity remains challenging because evolving a protease to recognize a new substrate, without counterselecting against its native substrate, often results in high residual activity on the original substrate. To address this, we developed Protease Engineering with Reactant Residence Time Control (PERRC), a platform that exploits the correlation between endoplasmic reticulum (ER) retention sequence strength and ER residence time. PERRC allows precise control over the stringency of protease evolution by adjusting counterselection to selection substrate ratios. Using PERRC, we evolved an orthogonal tobacco etch virus protease variant, TEVESNp, that selectively cleaves a substrate (ENLYFES) that differs by only one amino acid from its parent sequence (ENLYFQS). TEVESNp exhibits a remarkable 65-fold preference for the evolved substrate, marking the first example of an engineered orthogonal protease driven by such a slight difference in substrate recognition. Furthermore, TEVESNp functions as a competent protease for constructing orthogonal protein circuits in bacteria, and molecular dynamic simulations analysis reveals subtle yet functionally significant active site rearrangements. PERRC is a modular dual-substrate display system that facilitates precise engineering of protease specificity.

synthetic biology↗