bioRxiv Science⌕ Search

Biology subjects

Rafacz, D.

Publications and source records attributed to Rafacz, D..

2 recordsLinked to original sources

The impact of negative data sampling on antimicrobial peptide prediction

Antimicrobial peptides (AMPs) are a heterogeneous group of short polypeptides that target microorganisms but also viruses and cancer cells. Due to their lower selection for resistance compared to traditional antibiotics, AMPs have been attracting the ever-growing attention from researchers, including bioinformaticians. Machine learning represents the most cost-effective method for novel AMP discovery and consequently many computational tools for AMP prediction have been recently developed. In this article, we investigate the impact of negative data sampling on model performance and benchmarking. We generated 660 predictive models using 12 machine learning architectures, a single positive data set and 11 negative data sampling methods; the architectures and methods were defined on the basis of published AMP prediction software. Our results clearly indicate that similar training and benchmark data set, i.e. produced by the same or a similar negative data sampling method, positively affect model performance. Consequently, all the benchmark analyses that have been performed for AMP prediction models are significantly biased and, moreover, we do not know which model is the most accurate. To provide researchers with reliable information about the performance of AMP predictors, we also created a web server AMPBenchmark for fair model benchmarking. AMPBenchmark is available at http://BioGenies.info/AMPBenchmark.

bioinformatics↗

PCRedux: A Data Mining and Machine Learning Toolkit for qPCR Experiments

MotivationQuantitative Real-time PCR (qPCR) is a widely used -omics method for the precise quantification of nucleic acids, in which the result is associated with the presence/absence or quantity of a specific nucleic acid sequence. As the amount of qPCR data increases worldwide, the manual assessment of results becomes challenging and difficult to reproduce. To overcome this, some automatable characteristics of amplification curves have been described in the literature, often with an appropriate "rule of thumb". ResultsWe developed PCRedux to analyze and calculate 90 numerical qPCR amplification curve descriptors ( features") from large datasets of qPCR amplification curves that are aimed for interpretable machine learning and development of decision support systems. In a case study of a diverse dataset with 3181 positive, negative and ambiguous amplification curves, as assessed by three human raters, we demonstrate a sensitivity >99 % and specificity >97 % in detecting positive and negative amplification. PCRedux is unique as it goes beyond traditional qPCR analysis to capture curvature properties that improve the characterization and classification of amplification curves. The calculation of the features is reproducible and objective, since R is used as a controllable working environment. PCRedux is not a black box, but open source software following on the principle of mathematically interpretable features. These can be combined with user-defined labels for automatic multi-category classification and regression in machine learning. Availabilityhttps://cran.r-project.org/package=PCRedux. Web server: http://shtest.evrogen.net/PCRedux-app/. Documentation: https://PCRuniversum.github.io/PCRedux/.

bioinformatics↗