bioRxiv Science⌕ Search

Biology subjects

Hernandez-Salmeron, J. E.

Publications and source records attributed to Hernandez-Salmeron, J. E..

3 recordsLinked to original sources

Fast genome-based delimitation of Enterobacterales species

Average Nucleotide Identity (ANI) is becoming a standard measure for bacterial species delimitation. However, its calculation can take orders of magnitude longer than fast similarity estimates based on sampling of short nucleotides, compiled into so-called sketches. These estimates are widely used and correlate well with ANI. However, they might not be as accurate. Thus, we compared two sketching programs, mash and dashing, against ANI, in delimiting species among publicly available Esterobacterales genomes. Receiver Operating Characteristic (ROC) curve analysis found all three programs to be highly accurate, with Area Under the Curve (AUC) values of 0.99, indicating almost perfect species discrimination. Subsampling to reduce over-represented species, reduced these AUC values to 0.92. Focused tests with ten genera represented by more than three species, also showed almost identical results for all methods. Shigella showed the lowest AUC values (0.68), followed by Citrobacter (0.80). All other genera, Dickeya, Enterobacter, Escherichia, Klebsiella, Pectobacterium, Proteus, Providencia and Yersinia, produced AUC values above 0.90. The species delimitation thresholds varied, with species distance ranges in a few genera overlapping the genus ranges of other genera. Mash was able to separate the E. coli + Shigella complex into 25 apparent phylogroups. Testing mash for species separation in genera outside Enterobacterales showed AUCs above 0.95, again with different thresholds for species delimitation within each genus. Overall, our results suggest that fast estimates of genome similarity are as good as ANI for species delimitation. Therefore, these fast estimates might suffice for determining the role of genomic similarity in bacterial taxonomy.

bioinformatics↗

ANI, Mash and Dashing equally differentiate between Klebsiella species

Species of the genus Klebsiella are among the most important multidrug resistant human pathogens, though they have been isolated from a variety of environments. Given the need for quickly and accurately classifying newly sequenced Klebsiella genomes, we compared 982 Klebsiella genomes using different species-delimiting measures: Average Nucleotide Identity (ANI), which is becoming a standard for species delimitation, as well as Mash, Dashing, and DNA compositional signatures, which can be run in a fraction of the time required to run ANI. ROC analyses showed equal quality in species delimitation for ANI, Mash and Dashing (AUC: 0.99), followed by DNA signatures (AUC: 0.96). The groups obtained at optimal cutoffs were largely in agreement with species designation. Using optimized cutoffs, we obtained 17 species-level groups using either ANI, Mash, or Dashing, all containing the same genomes, unlike DNA signatures which broke the dataset into 38 groups. Further use of Mash to map species after adding draft genomes to the dataset also showed excellent results (AUC: 0.99), producing a total of 28 Klebsiella species in the publicly available genome collection. The ecological niches of Klebsiella strains were found to neither be related to species delimitation, nor to protein functional content, suggesting that a single Klebsiella species can have a wide repertoire of ecological functions.

genomics↗

Progress in quickly finding orthologs as reciprocal best hits

IntroductionFinding orthologs remains an important bottleneck in comparative genomics analyses. While the authors of software for the quick comparison of protein sequences evaluate the speed of their software and compare their results against the most usual software for the task, it is not common for them to evaluate their software for more particular uses, such as finding orthologs as reciprocal best hits (RBH). Here we compared RBH results, between prokaryotic genomes, obtained using software that runs faster than blastp. Namely, lastal, diamond, and MMseqs2. ResultsWe found that lastal required the least time to produce results. However, it yielded fewer results than any other program when comparing evolutionarily distant genomes. The program producing the most similar number of RBH as blastp was MMseqs2. This program also resulted in the lowest error estimates among the programs tested. The results with diamond were very close to those obtained with MMseqs2, with diamond running faster. Our results suggest that the best of the programs tested was diamond, ran with the "sensitive" option, which took 7% of the time as blastp to run, and produced results with lower error rates than blastp. AvailabilityA program to obtain reciprocal best hits using the software we tested is maintained at https://github.com/Computational-conSequences/SequenceTools

bioinformatics↗