bioRxiv ScienceSearch

Biology subjects

Kari, L.

Publications and source records attributed to Kari, L..

2 recordsLinked to original sources

ML-DSP: Machine Learning with Digital Signal Processing for ultrafast, accurate, and scalable genome classification at all taxonomic levels

BackgroundAlthough methods and software tools abound for the comparison, analysis, identification, and taxonomic classification of the enormous amount of genomic sequences that are continuously being produced, taxonomic classification remains challenging. The difficulty lies within both the magnitude of the dataset and the intrinsic problems associated with classification. The need exists for an approach and software tool that addresses the limitations of existing alignment-based methods, as well as the challenges of recently proposed alignment-free methods.\n\nResultsWe combine supervised Machine Learning with Digital Signal Processing to design ML-DSP, an alignment-free software tool for ultrafast, accurate, and scalable genome classification at all taxonomic levels.\n\nWe test ML-DSP by classifying 7,396 full mitochondrial genomes from the kingdom to genus levels, with 98% classification accuracy. Compared with the alignment-based classification tool MEGA7 (with sequences aligned with either MUSCLE, or CLUSTALW), ML-DSP has similar accuracy scores while being significantly faster on two small benchmark datasets (2,250 to 67,600 times faster for 41 mammalian mitochondrial genomes). ML-DSP also successfully scales to accurately classify a large dataset of 4,322 complete vertebrate mtDNA genomes, a task which MEGA7 with MUSCLE or CLUSTALW did not complete after several hours, and had to be terminated. ML-DSP also outperforms the alignment-free tool FFP (Feature Frequency Profiles) in terms of both accuracy and time, being three times faster for the vertebrate mtDNA genomes dataset.\n\nConclusionsWe provide empirical evidence that ML-DSP distinguishes complete genome sequences at all taxonomic levels. Ultrafast and accurate taxonomic classification of genomic sequences is predicted to be highly relevant in the classification of newly discovered organisms, in distinguishing genomic signatures, in identifying mechanistic determinants of genomic signatures, and in evaluating genome integrity.

genomics

An open-source k-mer based machine learning tool for fast and accurate subtyping of HIV-1 genomes

For many disease-causing virus species, global diversity is clustered into a taxonomy of subtypes with clinical significance. In particular, the classification of infections among the subtypes of human immunodeficiency virus type 1 (HIV-1) is a routine component of clinical management, and there are now many classification algorithms available for this purpose. Although several of these algorithms are similar in accuracy and speed, the majority are proprietary and require laboratories to transmit HIV-1 sequence data over the network to remote servers. This potentially exposes sensitive patient data to unauthorized access, and makes it impossible to determine how classifications are made and to maintain the data provenance of clinical bioinformatic workflows. We propose an open-source supervised and alignment-free subtyping method (KO_SCPCAPAMERISC_SCPCAP) that operates on k-mer frequencies in HIV-1 sequences. We performed a detailed study of the accuracy and performance of subtype classification in comparison to four state-of-the-art programs. Based on our testing data set of manually curated real-world HIV-1 sequences (n = 2, 784), Kameris obtained an overall accuracy of 97%, which matches or exceeds all other tested software, with a processing rate of over 1,500 sequences per second. Furthermore, our fully standalone general-purpose software provides key advantages in terms of data security and privacy, transparency and reproducibility. Finally, we show that our method is readily adaptable to subtype classification of other viruses including dengue, influenza A, and hepatitis B and C virus.

bioinformatics