bioRxiv Science⌕ Search

Biology subjects

Cemic, F.

Publications and source records attributed to Cemic, F..

2 recordsLinked to original sources

Generative models for antimicrobial peptide design: auto-encoders and beyond

BackgroundSince the number of multi-resistant pathogens is growing rapidly, new strategies to accelerate the development of antimicrobial drugs are urgently needed. A promising candidate class for new antibiotics are antimicrobial peptides, showing lower tendency to induce antibiotic resistance. High-throughput in silico strategies for candidate mining, such as generative deep learning algorithms, have become popular over the last few years and offer novel ways for peptide discovery. MethodsThis study presents a comparative analysis of contemporary deep learning models generative performance for generating novel antimicrobial peptides. The models examined include Variational Auto-Encoders, a Wasserstein Auto-Encoder, a Recurrent Neural Network and a Language Model. The primary focus of this study is the systematic comparison and evaluation of various methods and sampling options to identify the most suitable model and sampling strategy combination for different use cases. ResultsThe findings demonstrate the models capacity to generate peptide sequences exhibiting analogous properties to those of naturally occurring active peptides, which are utilized for model training while featuring an appropriate degree of sequence diversity. Auto-encoder-based models, particularly the Wasserstein auto-encoder, have generated novel and remarkably diverse sequences compared to recurrent neural networks and language models. This model category exhibits a propensity to prioritize the frequencies of individual amino acids during the learning process, in contrast to variational auto-encoders. Furthermore, latent space models have been shown to possess the capacity to utilize diverse methodologies for generating novel peptides. However, it is imperative to note that these sampling strategies are not universally advantageous or disadvantageous; their optimal selection is contingent on the specificities of each individual use case. ConclusionThe present study investigates the strengths and weaknesses of various generative models for antimicrobial peptides and suggests which model and sampling strategy combination should be favoured for specific individual applications.

bioinformatics↗

sORFdb - A database for sORFs, small proteins, and small protein families in bacteria

Small proteins with fewer than 100, particularly fewer than 50, amino acids are still largely unexplored. Nonetheless, they represent an essential part of bacterias often neglected genetic repertoire. In recent years, the development of ribosome profiling protocols has led to the detection of an increasing number of previously unknown small proteins. Despite this, they are overlooked in many cases by automated genome annotation pipelines, and often, no functional descriptions can be assigned due to a lack of known homologs. To understand and overcome these limitations, the current abundance of small proteins in existing databases was evaluated, and a new dedicated database for small proteins and their potential functions, called sORFdb, was created. To this end, small proteins were extracted from annotated bacterial genomes in the GenBank database. Subsequently, they were quality-filtered, compared, and complemented with proteins from Swiss-Prot, UniProt, and SmProt to ensure reliable identification and characterization of small proteins. Families of similar small proteins were created using bidirectional best BLAST hits followed by Markov clustering. Analysis of small proteins in public databases revealed that their number is still limited due to historical and technical constraints. Additionally, functional descriptions were often missing despite the presence of potential homologs. As expected, a taxonomic bias was evident in over-represented clinically relevant bacteria. This new and comprehensive database is accessible via a feature-rich website providing specialized search features for sORFs and small proteins of high quality. Additionally, small protein families with Hidden Markov Models and information on taxonomic distribution and other physicochemical properties are available. In conclusion, the novel small protein database sORFdb is a specialized, taxonomy-independent database that improves the findability and classification of sORFs, small proteins, and their functions in bacteria, thereby supporting their future detection and consistent annotation. All sORFdb data is freely accessible via https://sorfdb.computational.bio.

bioinformatics↗