bioRxiv Science⌕ Search

Biology subjects

Nayak, S. D.

Publications and source records attributed to Nayak, S. D..

2 recordsLinked to original sources

S9BactDB: A database for S9 family of proteases in bacterial genomes

S9 family proteins are serine proteases, divided into four subfamilies, involved in several functions associated with cell signalling, defence response and development. However, the annotation, characterization and statistical information are lacking. This is compounded by the huge number of bacterial genomes available. Hence, we have performed computational searches for S9 family peptidases as a step towards organising, curating the sequences and classifying into subfamilies. We have analysed annotated S9 family sequences from [~]32000 bacterial genomes/proteomes. All the curated information are presented as S9BacDB database (http://caps.ncbs.res.in/S9BactDB), provided in a user friendly way. The database provides various features such as information on the assemblies used (assembly, BioProject and BioSample details of each strain), the annotated POPs and their statistical distribution, unique domain architectures and the associated sequences and curated phylogenetic analysis. In addition, it provides the unique motifs and the associated information for Prolyl OligoPeptidases (POP) subtypes. An ML model is also integrated that can classify recognised sequences into a S9 sub family category. In conclusion, S9BactDb is a comprehensive platform that provides meticulously curated data on statistical/bioinformatics analysis of all S9 family proteins originating from fully sequenced bacterial genomes in the RefSeq database.

bioinformatics↗

Identification and study of Prolyl Oligopeptidases and related sequences in bacterial lineages

Proteases are enzymes that break down proteins, and serine proteases are an important subset of these enzymes. Prolyl oligopeptidase (POP) is a family of serine proteases (S9 family) that has the ability to cleave peptide bonds involving proline residues and it is unique for its ability to cleave various small oligopeptides shorter than 30 amino acids. The S9 family from the MEROPS database, is classified into four subfamilies based on active site motifs. These S9 subfamilies assume a crucial position owing to their diverse biological roles and potential therapeutic applications in various diseases. In this study, we have examined [~]32000 completely annotated bacterial genomes from the NCBI RefSeq Assembly database to identify annotated S9 family proteins. This results in the discovery of [~]53,000 bacterial S9 family proteins (referred to as POP homologues). These sequences are classified into distinct subfamilies through various machine-learning approaches and comprehensive analysis of their distribution across various phyla and species and domain architecture analysis are also conducted. Distinct subclusters and class-specific motifs of POPs were identified, suggesting differences in substrate specificity in POP homologues. This study can enable future research of these gene families that are involved in many important biological processes.

bioinformatics↗