bioRxiv Science⌕ Search

Biology subjects

Savojardo, C.

Publications and source records attributed to Savojardo, C..

3 recordsLinked to original sources

CoCoNat: a novel method based on deep-learning for coiled-coil prediction

MotivationCoiled-coil domains (CCD) are widespread in all organisms performing several crucial functions. Given their relevance, the computational detection of coiled-coil domains is very important for protein functional annotation. State-of-the art prediction methods include the precise identification of coiled-coil domain boundaries, the annotation of the typical heptad repeat pattern along the coiled-coil helices as well as the prediction of the oligomerization state. ResultsIn this paper we describe CoCoNat, a novel method for predicting coiled-coil helix boundaries, residue-level register annotation and oligomerization state. Our method encodes sequences with the combination of two state-of-the-art protein language models and implements a three-step deep learning procedure concatenated with a Grammatical-Restrained Hidden Conditional Random Field (GRHCRF) for CCD identification and refinement. A final neural network (NN) predicts the oligomerization state. When tested on a blind test set routinely adopted, CoCoNat obtains a performance superior to the current state-of-the-art both for residue-level and segment-level coiled-coil detection. CoCoNat significantly outperforms the most recent state-of-the art method on register annotation and prediction of oligomerization states. AvailabilityCoCoNat is available at https://coconat.biocomp.unibo.it. Contactpierluigi.martelli@unibo.it

bioinformatics↗

ISPRED-SEQ: Deep neural networks and embeddings for predicting interaction sites in protein sequences.

The knowledge of protein-protein interaction sites (PPIs) is crucial for protein functional annotation. Here we address the problem focusing on the prediction of putative PPIs having as input protein sequences. The problem is important given the huge volume of sequences compared to experimental and/or computed protein structures. Taking advantage of recently developed protein language models and Deep Neural networks here we describe ISPRED-SEQ, which overpasses state-of-the-art predictors addressing the same problem. ISPRED-SEQ is freely available for testing at https://ispredws.biocomp.unibo.it.

bioinformatics↗

E-SNPs&GO: Embedding of protein sequence and function improves the annotation of pathogenic variants.

MotivationThe advent of massive DNA sequencing technologies is producing a huge number of human single-nucleotide polymorphisms occurring in protein-coding regions and possibly changing protein sequences. Discriminating harmful protein variations from neutral ones is one of the crucial challenges in precision medicine. Computational tools based on artificial intelligence provide models for protein sequence encoding, bypassing database searches for evolutionary information. We leverage the new encoding schemes for an efficient annotation of protein variants. ResultsE-SNPs&GO is a novel method that, given an input protein sequence and a single residue variation, can predict whether the variation is related to diseases or not. The proposed method, for the first time, adopts an input encoding completely based on protein language models and embedding techniques, specifically devised to encode protein sequences and GO functional annotations. We trained our model on a newly generated dataset of 65,888 human protein single residue variants derived from public resources. When tested on a blind set comprising 6,541 variants, our method outperforms recent approaches released in literature for the same task, reaching a MCC score of 0.71. We propose E-SNPs&GO as a suitable, efficient and accurate large-scale annotator of protein variant datasets. Contactpierluigi.martelli@unibo.it

bioinformatics↗