bioRxiv ScienceSearch

Biology subjects

Blackwell, G. A.

Publications and source records attributed to Blackwell, G. A..

2 recordsLinked to original sources

Exploring bacterial diversity via a curated and searchable snapshot of archived DNA sequences

The open sharing of genomic data provides an incredibly rich resource for the study of bacterial evolution and function, and even anthropogenic activities such as the widespread use of antimicrobials. Whilst these archives are rich in data, considerable processing is required before biological questions can be addressed. Here, we assembled and characterised 661,405 bacterial genomes using a uniform standardised approach, retrieved from the European Nucleotide Archive (ENA) in November of 2018. A searchable COBS index has been produced, facilitating the easy interrogation of the entire dataset for a specific gene or mutation. Additional MinHash and pp-sketch indices support genome-wide comparisons and estimations of genomic distance. An analysis on this scale revealed the uneven species composition in the ENA/public databases, with just 20 of the total 2,336 species making up 90% of the genomes. The over-represented species tend to be acute/common human pathogens. This aligns with research priorities at different levels from individuals with targeted but focused research questions, areas of focus for the funding bodies or national public health agencies, to those identified globally as priority pathogens by the WHO for their resistance to front and last line antimicrobials. Understanding the actual and potential biases in bacterial diversity depicted in this snapshot, and hence within the data being submitted to the public sequencing archives, is essential if we are to target and fill gaps in our understanding of the bacterial kingdom.

microbiology

gbpA and chiA genes are not uniformly distributed amongst diverse Vibrio cholerae

Members of the bacterial genus Vibrio utilise chitin both as a metabolic substrate and a signal to activate natural competence. Vibrio cholerae is a bacterial enteric pathogen, sub-lineages of which can cause pandemic cholera. However, the chitin metabolic pathway in V. cholerae has been dissected using only a limited number of laboratory strains of this species. Here, we survey the complement of key chitin metabolism genes amongst 195 diverse V. cholerae. We show that the gene encoding GbpA, known to be an important colonisation and virulence factor in pandemic isolates, is not ubiquitous amongst V. cholerae. We also identify a putatively novel chitinase, and present experimental evidence in support of its functionality. Our data indicate that the chitin metabolic pathway within the V. cholerae species is more complex than previously thought, and emphasise the importance of considering genes and functions in the context of a species in its entirety, rather than simply relying on traditional reference strains. Impact statementIt is thought that the ability to metabolise chitin is ubiquitous amongst Vibrio spp., and that this enables these species to survive in aqueous and estuarine environmental contexts. Although chitin metabolism pathways have been detailed in several members of this genus, little is known about how these processes vary within a single Vibrio species. Here, we present the distribution of genes encoding key chitinase and chitin-binding proteins across diverse Vibrio cholerae, and show that our canonical understanding of this pathway in this species is challenged when isolates from non-pandemic V. cholerae lineages are considered alongside those linked to pandemics. Furthermore, we show that genes previously thought to be species core genes are not in fact ubiquitous, and we identify novel components of the chitin metabolic cascade in this species, and present functional validation for these observations. Data summaryThe authors confirm that all supporting data, code, and protocols have been provided within the article or through supplementary data files. O_LINo whole-genome sequencing data were generated in this study. Accession numbers for the publicly-available sequences used for these analyses are listed in Supplementary Table 1, Table 2, and the Methods. C_LIO_LIAll other data which underpin the figures in this manuscript, including pangenome data matrices, modified and unmodified sequence alignments and phylogenetic trees, original images of gels and immunoblots, raw fluorescence data, amplicon sequencing reads, and the R code used to generate Figure 7, are available in Figshare: https://dx.doi.org/10.6084/m9.figshare.13169189 (Note for peer-review: Figshare DOI is inactive but will be activated upon publication, please use temporary URL https://figshare.com/s/7795a2d80c13f694f8fa for review). C_LI O_FIG O_LINKSMALLFIG WIDTH=183 HEIGHT=200 SRC="FIGDIR/small/430729v1_fig7.gif" ALT="Figure 7"> View larger version (38K): org.highwire.dtl.DTLVardef@100c246org.highwire.dtl.DTLVardef@d295aborg.highwire.dtl.DTLVardef@1601f69org.highwire.dtl.DTLVardef@1ae3ecb_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 7.C_FLOATNO ChiA-3-6xHis displays chitobiosidase and endochitinase activities, but not -N-acetylglucosaminidase activity. Lysates and supernatants included in Figures 6b and 6c were assayed for chitinase enzyme activity using a fluorometric chitinase assay kit (see Methods for details). Lysed cells from E. coli cultures harbouring pMJD157 and cultured in the presence of arabinose were the only samples which produced detectable and statistically significant signals on triacetylchitotriose and chitobiose substrates (a, b). No signal was detected in the presence of glucosaminide substrate (c). All plots are scaled equivalently. P = pellet; S = supernatant; EV = empty vector (pBAD33). Parametric t-tests performed where indicated: ns = not significant; *** = p<0.001; **** = p<0.0001. C_FIG

microbiology