bioRxiv Science⌕ Search

Biology subjects

Sanchez, J. E.

Publications and source records attributed to Sanchez, J. E..

3 recordsLinked to original sources

JRSeek: Artificial Intelligence Meets Jelly Roll Fold Classification in Viruses

The jelly roll (JR) fold is the most common structural motif found in the capsid and nucleocapsid of viruses. Its pervasiveness across many different viral families motives developing a tool to predict its presence from a sequence. In the current work, logistic regression (LR) models trained on six different large language model (LLM) embeddings exhibited over 95% accuracy in differentiating JR from non-JR sequences. The dataset used for training and testing included sequences from single JR viruses, non-JR viruses, and non-virus immunoglobulin-like {beta}-sandwich (IGLBS) proteins which closely resemble the JR fold in structure. The high accuracy is particularly remarkable given the low sequence similarity across viral families and the balanced nature of the dataset. Also, the accuracy of the models was independent of LLM embeddings, suggesting that peak accuracy for predicting viral JR folds hinges more on the data quality and quantity rather than on the specific mathematical models used. Given that many viral capsid and nucleocapsid structures have yet to be resolved, using sequence-based LLMs is a promising strategy that can readily be applied to available data. Principal Component Analysis of the Bert-U100 embeddings demonstrates that most IGLBS sequences and a subset of JR and non-JR sequences are distinguishable even before the application of the LR model, but the LR model is necessary to differentiate a subset of more ambiguous sequences. When applied to double JR folds, the Bert-U100 model was able to assign the JR motif for some viral families, providing evidence for the models generalizability. However, for other families, this generalizability was not observed, motivating a future need to develop other models informed by double JR folds. Lastly, the Bert-U100 model was also able to predict whether sequences from a dataset of unclassified viruses produce the JR fold. Two examples are given and the JR predictions are corroborated by AlphaFold3. Altogether, this work demonstrates that JR folds can, in principle, be predicted from their sequences.

bioinformatics↗

Local Microenvironments of capsomer variants in the PBCV-1

PBCV-1, a giant virus classified among the Nucleocytoviricota virus (NCV) whose structure has been determined to near atomic resolution. The majority capsomers forming the capsid of PBCV-1 are Type I capsomers while five other type of variants have been found in recent high resolution structure. Interestingly, some variants, such as Type V capsomers, are found at particular capsid locations whose roles are unclear. To reveal the roles of a Type V capsomer, we replaced the Type V capsomer by a Type I capsomer to compare the interaction among the two types of capsomer variant, especially the interactions between each of the Type V/Type I capsomer and its local capsid microenvironment. Our results revealed significant differences between Type V and Type I capsomers. Notably, the Type V capsomer demonstrated a stronger binding force to the surrounding capsomers than the Type I capsomer. Moreover, the identified salt bridges between Type V/I capsomers and their surrounding capsomers corroborate the results of electrostatic calculations, further highlighting the important residues involved in these interactions. Understanding these local capsid microenvironments will be essential to elucidate the mechanisms governing viral capsid assembly.

biophysics↗

BioBrigit, A Hybrid Deep Learning and Knowledge-based Approach to Model Metal Pathways in Proteins: Application to a Di-Copper Tyrosinase

The interaction of metallic species with proteins has been fundamental in evolution and key in many physiological processes. How metals bind to proteins also holds promise in many fields, like the design of new biocatalysts or the fight against pathogens. Nonetheless, uncovering the mechanism under which proteins recruit metal ions is far from understood and is one of the challenges in bioinorganic chemistry and structural biology. Computational methods are potentially among the most promising tools for this endeavor. Only a handful of efficient structural predictors of metal binding sites exist to date. Most focus on identifying the most stable binding sites in the protein scaffolds. Although these methods are very interesting, they do not consider the exploration of transient, sub-optimal binding sites that could be relevant in metal binding pathways in proteins. At the far end of modeling capabilities nowadays, we introduce BioBrigit, a hybrid Deep Learning - knowledge-based approach that suggests metal binding pathways in proteins. To demonstrate the methods viability, we apply it to the di-copper tyrosinase from Streptomyces castaneoglobisporus, a system for which crystallographic experiments allowed the identification of a series of transient sites of the copper in its path from a chaperone to the final catalytic site. Combined with homology modeling and large-scale molecular dynamics, BioBrigit allows for computational characterization of all experimental sites and for better understanding of the copper recruitment mechanism. BioBrigit appears as an asset in a field full of unknowns like metal binding to proteins and opens the way to further algorithms in this area. Source code, documentation, and data are available at https://github.com/insilichem/BioBrigit

bioinformatics↗