bioRxiv Science⌕ Search

Biology subjects

Febrer Martinez, P.

Publications and source records attributed to Febrer Martinez, P..

2 recordsLinked to original sources

CryptoBank: A Resource for the Identification and Prediction of Cryptic Sites in Proteins

Cryptic binding sites in proteins, which are hidden in the absence of a ligand, offer opportunities to modulate targets previously considered undruggable. However, the scarcity of experimentally validated examples limits the development of predictive tools. Here, we introduce CryptoBank, a large-scale database of cryptic sites identified by applying a machine learning model to detect ligand-induced conformational changes in over 5.5 million structural alignments of unbound (apo) and bound (holo) protein pairs from the Protein Data Bank (PDB). Our analysis reveals that cryptic pockets are widespread, occurring in approximately 16.3% of protein clusters. Leveraging this resource, we fine-tuned a protein language model (PLM) to predict cryptic sites directly from protein sequence information. This sequence-based model achieves high precision (PR AUC 0.8) when query sequences share more than 20% identity with CryptoBank entries. Critically, we demonstrate its broader utility by predicting a cryptic site in human TPP1, a protein with less than 20% sequence identity to any CryptoBank entry, and validating its opening using molecular dynamics simulations. CryptoBank and the predictive PLM are publicly accessible via a web server, providing valuable resources for cryptic site discovery and drug development.

bioinformatics↗

Host-Guest binding free energies a la carte: an automated OneOPES protocol

Estimating absolute binding free energies from molecular simulations is a key step in computer-aided drug design pipelines, but agreement between computational results and experiments is still very inconsistent. Both the accuracy of the computational model and the quality of the statistical sampling contribute to this discrepancy, yet disentangling the two remains a challenge. In this study, we present an automated protocol based on OneOPES, an enhanced sampling method that exploits replica exchange and can accelerate several collective variables, to address the sampling problem. We apply this protocol to 37 host-guest systems. The simplicity of setting up the simulations and of producing well-converged binding free energy estimates without the need to optimize simulation parameters provides a reliable solution to the sampling problem. This, in turn, allows for a systematic force field comparison and ranking according to the correlation between simulations and experiments, which can inform the selection of an appropriate model. The protocol can be readily adapted to test more force field combinations and study more complex protein-ligand systems, where the choice of an appropriate physical model is often based on heuristic considerations rather than a systematic optimization.

biochemistry↗