bioRxiv Science⌕ Search

bioRxiv · 10.1101/2021.11.03.467004

Discordance between different bioinformatic methods for identifying resistance genes from short-read genomic data, with a focus on Escherichia coli

Abstract

2.Several bioinformatics genotyping algorithms are now commonly used to characterise antimicrobial resistance (AMR) gene profiles in whole genome sequencing (WGS) data, with a view to understanding AMR epidemiology and developing resistance prediction workflows using WGS in clinical settings. Accurately evaluating AMR in Enterobacterales, particularly Escherichia coli, is of major importance, because this is a common pathogen. However, robust comparisons of different genotyping approaches on relevant simulated and large real-life WGS datasets are lacking. Here, we used both simulated datasets and a large set of real E. coli WGS data (n=1818 isolates) to systematically investigate genotyping methods in greater detail. Simulated constructs and real sequences were processed using four different bioinformatic programs (ABRicate, ARIBA, KmerResistance, and SRST2, run with the ResFinder database) and their outputs compared. For simulations tests where 3,079 AMR gene variants were inserted into random sequence constructs, KmerResistance was correct for 3,076 (99.9%) simulations, ABRicate for 3,054 (99.2%), ARIBA for 2,783 (90.4%) and SRST2 for 2,108 (68.5%). For simulations tests where two closely related gene variants were inserted into random sequence constructs, KmerResistance the correct alleles in 35,338/46,318 (76.3%) ABRicate identified in 11,842/46,318 (25.6%) of simulations, ARIBA in 1679/46,318 (3.6%), and SRST2 in 2000/46,318 (4.3%). In real data, across all methods, 1392/1818 (76%) isolates had discrepant allele calls for at least one gene. Our evaluations revealed poor performance in scenarios that would be expected to be challenging (e.g. identification of AMR genes at <10x coverage, discriminating between closely related AMR gene sequences), but also identified systematic sequence classification (i.e. naming) errors even in straightforward circumstances, which contributed to 1081/3092 (35%) errors in our most simple simulations and at least 2530/4321 (59%) discrepancies in real data. Further, many of the remaining discrepancies were likely "artefactual" with reporting cut-off differences accounting for at least 1430/4321 (33%) discrepants. Comparing outputs generated by running multiple algorithms on the same dataset can help identify and resolve these artefacts, but ideally new and more robust genotyping algorithms are needed. 3. Impact statementWhole-genome sequencing is widely used for studying the epidemiology of antimicrobial resistance (AMR) genes in bacteria; however, there is some concern that outputs are highly dependent on the bioinformatics methods used. This work evaluates these concerns in detail by comparing four different, commonly used AMR gene typing methods using large simulated and real datasets. The results highlight performance issues for most methods in at least one of several simulated and real-life scenarios. However most discrepancies between methods were due to differential labelling of the same sequences related to the assumptions made regarding the underlying structure of the reference resistance gene database used (i.e. that resistance genes can be easily classified in well-defined groups). This study represents a major advance in quantifying and evaluating the nature of discrepancies between outputs of different AMR typing algorithms, with relevance for historic and future work using these algorithms. Some of the discrepancies can be resolved by choosing methods with fewer assumptions about the reference AMR gene database and manually resolving outputs generated using multiple programs. However, ideally new and better methods are needed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Davies, T. J., Swann, J., Sheppard, A. E., Pickford, H., Lipworth, S., AbuOun, M., Ellington, M., Fowler, P. W., Hopkins, S., Hopkins, K., Crook, D., Peto, T. E., Anjum, M. F., Walker, A. S., Stoesser, N.. 2021-11-03. Discordance between different bioinformatic methods for identifying resistance genes from short-read genomic data, with a focus on Escherichia coli. https://doi.org/10.1101/2021.11.03.467004

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A conserved cysteine-histidine-glutamate metal site identifies DUF501 (Rv1025), an essential uncharacterised protein family of Mycobacterium tuberculosis, as a candidate metalloenzyme and drug target

A substantial fraction of the Mycobacterium tuberculosis proteome remains functionally uncharacterised. Rv1025, a 155-residue protein carrying the domain of unknown function DUF501 (Pfam PF04417), is essential by transposon mutagenesis and vulnerable by CRISPR interference, an attractive but neglected drug target, yet has never been functionally described. The family (4,370 proteins, no Gene Ontology term, no solved structure) is uncharacterised across all organisms and essential in three Actinobacterial genera. A Foldseek search of the AlphaFold model against complete structural databases finds no significant homolog, indicating a novel fold. The operon eno-divIC-Rv1025-ppx2 is conserved across the Actinobacteria phylum, yet AlphaFold-Multimer finds no direct complex between Rv1025 and its neighbour DivIC. Instead, conservation across 8,700 homologous sequences reveals a near-invariant Cys113-His115-Glu59 cluster forming a pocket. Holo AlphaFold3 predictions with Zn, Fe and Mn confidently place a divalent metal on this triad at 2.25-2.47 A; mutating the triad relocates the metal, and an independent backbone-geometry predictor recovers the same site, confirming specificity. The triad is universal across the family: present in all 1,472 near-complete bacterial sequences of the Pfam alignment, with no non-conservative substitution among the 2,228 sequences examined, a defining feature of bacterial DUF501 rather than a mycobacterial peculiarity. We propose that DUF501 is a metal-binding protein and candidate metalloenzyme, the first functional hypothesis for this family, whose conserved, essential metal pocket is a promising drug target. As the predictions build on a conservation-defined site within a fully computational study, they are supportive rather than proof of metal occupancy and warrant experimental validation.

microbiology↗

Mycoplasmal endosymbionts of Trichomonas vaginalis are associated with reduced risk for Chlamydia trachomatis endometrial infection in asymptomatic, coinfected, women.

Trichomonas vaginalis is a protozoan parasite that causes trichomoniasis, the most common curable non-viral sexually transmitted infection, and Chlamydia trachomatis is a bacterial pathogen that can ascend to the upper genital tract and cause pelvic inflammatory disease, infertility, and ectopic pregnancy. T. vaginalis harbors bacterial endosymbionts, including Candidatus Malacoplasma girerdii, an obligate symbiont, and Metamycoplasma hominis, which can live freely or symbiotically. In a 16S rRNA sequencing study of the cervicovaginal microbiome of women at high risk for chlamydial infection, Ca. M. girerdii abundance was one of 13 features predicting lack of chlamydial spread to the endometrium, despite no direct association between T. vaginalis infection and reduced chlamydial ascension. Investigating the relationship between these microorganisms further, we found that T. vaginalis vaginal abundance correlated positively with chlamydial burden in women whose infection was confined to the cervix, while a nonsignificant inverse relationship was seen in women with endometrial spread. Among participants with high chlamydial burden, Ca. M. girerdii was detected exclusively in women without endometrial infection. Both endosymbionts trended toward more frequent detection, and higher abundance, in coinfected women without endometrial spread, while M. hominis abundance correlated strongly with T. vaginalis burden in this group. These findings suggest that mycoplasmal endosymbionts of T. vaginalis, rather than T. vaginalis itself, are microbial factors limiting chlamydial ascension, and point to a three-way interaction between parasite, endosymbiont, and bacterial pathogen that shapes upper genital tract C. trachomatis infection risk.

microbiology↗

Understanding the physiological alterations of Vibrio cholerae upon exposure to L-ascorbic acid

The scourge of cholera remains a major global public health threat. It affects up to 4 million people worldwide and causes tens of thousands of deaths each year. The disease is experiencing a concerning resurgence in many parts of Africa, the Middle East, and Asia. To effectively tackle cholera and circumvent rising antimicrobial resistance, targeted biological and preventive approaches, complementing traditional rehydration, are urgently needed. In this regard, our group has demonstrated the efficacy of L-ascorbic acid in controlling the growth and pathogenesis of Vibrio cholerae in vitro. The present work further provides a mechanistic elucidation of the L-ascorbic acid-mediated physiological changes in V. cholerae and also bolsters such a non-antibiotic approach to control cholera.

microbiology↗