bioRxiv Science⌕ Search

Biology subjects

Plantinga, N. L.

Publications and source records attributed to Plantinga, N. L..

2 recordsLinked to original sources

PlasmidEC and gplas2: An optimised short-read approach to predict and reconstruct antibiotic resistance plasmids in Escherichia coli

Accurate reconstruction of Escherichia coli antibiotic resistance gene (ARG) plasmids from Illumina sequencing data has proven to be a challenge with current bioinformatic tools. In this work, we present an improved method to reconstruct E. coli plasmids using short reads. We developed plasmidEC, an ensemble classifier that identifies plasmid-derived contigs by combining the output of three different binary classification tools. We showed that plasmidEC is especially suited to classify contigs derived from ARG plasmids with a high recall of 0.941. Additionally, we optimised gplas, a graph-based tool that bins plasmid-predicted contigs into distinct plasmid predictions. Gplas2 is more effective at recovering plasmids with large sequencing coverage variations and can be combined with the output of any binary classifier. The combination of plasmidEC with gplas2 showed a high completeness (median=0.818) and F1-score (median=0.812) when reconstructing ARG plasmids and exceeded the binning capacity of the reference-based method MOB-suite. In the absence of long read data, our method offers an excellent alternative to reconstruct ARG plasmids in E. coli. Data SummaryNo new sequencing data have been generated in this study. All genomes used in this research are publicly available at the GenBank and Sequence Read Archive of the National Center for Biotechnology Information. Accession numbers are specified in Supplementary Materials. Scripts to reproduce the results reported in this manuscript can be accessed at https://gitlab.com/jpaganini/ecoli-binary-classifier. The ensemble classifier, plasmidEC, is publicly available at https://gitlab.com/mmb-umcu/plasmidEC (release 1.3.1), and gplas2 (release 1.0.0) can be found at https://gitlab.com/mmb-umcu/gplas2. Impact StatementEscherichia coli has emerged as a highly pervasive multidrug resistant pathogen on a global scale. The dissemination of resistance is significantly influenced by plasmids, mobile genetic elements that facilitate the transfer of antimicrobial resistance genes within and between diverse bacterial species. Consequently, precise and high-throughput identification of plasmids is imperative for effective genomic surveillance of resistance. However, accurate plasmid reconstruction remains challenging with the use of affordable short-read sequencing data. In this work, we present a novel method to accurately predict and reconstruct E. coli plasmids based on Illumina data. Additionally, we demonstrate that our approach outperforms the reference-based method MOB-suite, especially when reconstructing plasmids carrying antimicrobial resistance genes.

bioinformatics↗

Recovering Escherichia coli plasmids in the absence of long-read sequencing data

The incidence of infections caused by multidrug-resistant E. coli strains has risen in the past years. Antibiotic resistance in E. coli is often mediated by acquisition and maintenance of plasmids. The study of E. coli plasmid epidemiology and genomics often requires long-read sequencing information, but recently a number of tools that allow plasmid prediction from short-read data have been developed. Here, we reviewed 25 available plasmid prediction tools and categorized them into binary plasmid/chromosome classification tools and plasmid reconstruction tools. We benchmarked six tools that aim to reliably reconstruct distinct plasmids, with a special focus on plasmids carrying antibiotic resistance genes (ARGs) such as extended-spectrum beta-lactamase genes. They use either assembly graph information (plasmidSPAdes, gplas), reference databases (MOB-Suite, FishingForPlasmids) or both (HyAsP and SCAPP) to produce plasmid predictions. The benchmark data set consisted of 240 E. coli strains, harboring 631 plasmids, which were representative for the diversity of E. coli in public databases. Notably, these strains were not used for training any of the tools. We found that two thirds (n=425, 66.3.%) of all plasmids were correctly reconstructed by at least one of the six tools, with a range of 92 (14.58%) to 317 (50.23%) correctly predicted plasmids. However, the majority of plasmids that carried antibiotic resistance genes (n=85, 57.8%) could not be completely recovered as distinct plasmids by any of the tools. MOB-suite was the only tool that was able to correctly reconstruct the majority of plasmids (n=317, 50.23%), and performed best at reconstructing large plasmids (n=166, 46.37%) and ARG-plasmids (n=41, 27.9%), but predictions frequently contained chromosome contamination (40%). In contrast, plasmidSPAdes reconstructed the highest fraction of plasmids smaller than 18 kbp (n=168, 61.54%). Large ARG-plasmids, however, were recovered with small precision values (median=0.47, IQR=0.61), indicating that plasmidSPAdes frequently merged sequences derived from distinct replicons. Additionally, only 63% of all plasmid-borne ARGs were correctly predicted by plasmidSPAdes. The remaining four tools (FishingForPlasmids, HyAsP, SCAPP and gplas) were able to correctly reconstruct a combined total of 18 plasmids that were missed by MOB-suite and plasmidSPAdes. Available bioinformatic tools can provide valuable insight into E. coli plasmids, but also have important limitations. This work will serve as a guideline for selecting the most appropriate plasmid reconstruction tool for studies focusing on E. coli plasmids in the absence of long-read sequencing data.

bioinformatics↗