bioRxiv ScienceSearch

Biology subjects

Heider, D.

Publications and source records attributed to Heider, D..

2 recordsLinked to original sources

γBOriS: Identification of Origins of Replication in Gammaproteobacteria using Motif-based Machine Learning

The biology of bacterial cells is, in general, based on the information encoded on circular chromosomes. Regulation of chromosome replication is an essential process which mostly takes place at the origin of replication (oriC). Identification of high numbers of oriC is a prerequisite to enable systematic studies that could lead to insights of oriC functioning as well as novel drug targets for antibiotic development. Current methods for identyfing oriC sequences rely on chromosome-wide nucleotide disparities and are therefore limited to fully sequenced genomes, leaving a superabundance of genomic fragments unstudied. Here, we present{gamma} BOriS (Gammaproteobacterial oriC Searcher), which accurately identifies oriC sequences on gammaproteobacterial chromosomal fragments by employing motif-based DNA classification. Using{gamma} BOriS, we created BOriS DB, which currently contains 25,827 oriC sequences from 1,217 species, thus making it the largest available database for oriC sequences to date.

bioinformatics

Deep Learning on Chaos Game Representation for Proteins

Classification of protein sequences is one big task in bioinformatics and has many applications. Different machine learning methods exist and are applied on these problems, such as support vector machines (SVM), random forests (RF), and neural networks (NN). All of these methods have in common that protein sequences have to be made machine-readable and comparable in the first step, for which different encodings exist. These encodings are typically based on physical or chemical properties of the sequence. However, due to the outstanding performance of deep neural networks (DNN) on image recognition, we used frequency matrix chaos game representation (FCGR) for encoding of protein sequences into images. In this study, we compare the performance of SVMs, RFs, and DNNs, trained on FCGR encoded protein sequences. While the original chaos game representation (CGR) has been used mainly for genome sequence encoding and classification, we modified it to work also for protein sequences, resulting in n-flakes representation, an image with several icosagons.\n\nWe could show that all applied machine learning techniques (RF, SVM, and DNN) show promising results compared to the state-of-the-art methods on our benchmark datasets, with DNNs outperforming the other methods and that FCGR is a promising new encoding method for protein sequences.

bioinformatics