bioRxiv ScienceSearch

Biology subjects

Page, A. J.

Publications and source records attributed to Page, A. J..

5 recordsLinked to original sources

Candidatus Ornithobacterium hominis sp. nov.: insights gained from draft genomes obtained from nasopharyngeal swabs

Candidatus Ornithobacterium hominis sp. nov. represents a new member of the Flavobacteriaceae detected in 16S rRNA gene surveys from Southeast Asia, Africa and Australia. It frequently colonises the infant nasopharynx at high proportional abundance, and we demonstrate its presence in 42% of nasopharyngeal swabs from 12 month old children in the Maela refugee camp in Thailand. The species, a Gram negative bacillus, has not yet been cultured but the cells can be identified in mixed samples by fluorescent hybridisation. Here we report seven genomes assembled from metagenomic data, two to improved draft standard. The genomes are approximately 1.9Mb, sharing 62% average amino acid identity with the only other member of the genus, the bird pathogen Ornithobacterium rhinotracheale. The draft genomes encode multiple antibiotic resistance genes, competition factors, Flavobacterium johnsoniae-like gliding motility genes and a homolog of the Pasteurella multocida mitogenic toxin. Intra- and inter-host genome comparison suggests that colonisation with this bacterium is both persistent and strain exclusive.

genomics

PlasmidTron: assembling the cause of phenotypes from NGS data

When defining bacterial populations through whole genome sequencing (WGS) the samples often have detailed associated metadata that relate to disease severity, antimicrobial resistance, or even rare biochemical traits. When comparing these bacterial populations, it is apparent that some of these phenotypes do not follow the phylogeny of the host i.e. they are genetically unlinked to the evolutionary history of the host bacterium. One possible explanation for this phenomenon is that the genes are moving independently between hosts and are likely associated with mobile genetic elements (MGE). However, identifying the element that is associated with these traits can be complex if the starting point is short read WGS data. With the increased use of next generation WGS in routine diagnostics, surveillance and epidemiology a vast amount of short read data is available and these types of associations are relatively unexplored. One way to address this would be to perform assembly de novo of the whole genome read data, including its MGEs. However, MGEs are often full of repeats and can lead to fragmented consensus sequences. Deciding which sequence is part of the chromosome, and which is part of a MGE can be ambiguous. We present PlasmidTron, which utilises the phenotypic data normally available in bacterial population studies, such as antibiograms, virulence factors, or geographic information, to identify sequences that are likely to represent MGEs linked to the phenotype. Given a set of reads, categorised into cases (showing the phenotype) and controls (phylogenetically related but phenotypically negative), PlasmidTron can be used to assemble de novo reads from each sample linked by a phenotype. A k-mer based analysis is performed to identify reads associated with a phylogenetically unlinked phenotype. These reads are then assembled de novo to produce contigs. By utilising k-mers and only assembling a fraction of the raw reads, the method is fast and scalable to large datasets. This approach has been tested on plasmids, because of their contribution to important pathogen associated traits, such as AMR, hence the name, but there is no reason why this approach cannot be utilized for any MGE that can move independently through a bacterial population. PlasmidTron is written in Python 3 and available under the open source licence GNU GPL3 from https://github.com/sanger-pathogens/plasmidtron.\n\nDATA SUMMARYO_LISource code for PlasmidTron is available from Github under the open source licence GNU GPL 3; (url - https://goo.gl/ot6rT5)\nC_LIO_LISimulated raw reads files have been deposited in Figshare; (url - https://doi.org/10.6084/m9.figshare.5406355.vl)\nC_LIO_LISalmonella enterica serovar Weltevreden strain VNS10259 is available from GenBank; accession number GCA_001409135.\nC_LIO_LISalmonella enterica serovar Typhi strain BL60006 is available from GenBank; accession number GCA_900185485.\nC_LIO_LIAccession numbers for all of the Illumina datasets used in this paper are listed in the supplementary tables.\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTPlasmidTron utilises the phenotypic data normally available in bacterial population studies, such as antibiograms, virulence factors, or geographic information, to identify sequences that are likely to represent MGEs linked to the phenotype.

bioinformatics

SeroBA: rapid high-throughput serotyping of Streptococcus pneumoniae from whole genome sequence data

Streptococcus pneumoniae is responsible for 240,000 - 460,000 deaths in children under 5 years of age each year. Accurate identification of pneumococcal serotypes is important for tracking the distribution and evolution of serotypes following the introduction of effective vaccines. Recent efforts have been made to infer serotypes directly from genomic data but current software approaches are limited and do not scale well. Here, we introduce a novel method, SeroBA, which uses a hybrid assembly and mapping approach. We compared SeroBA against real and simulated data and present results on the concordance and computational performance against a validation dataset, the robustness and scalability when analysing a large dataset, and the impact of varying the depth of coverage in the cps locus region on sequence-based serotyping. SeroBA can predict serotypes, by identifying the cps locus, directly from raw whole genome sequencing read data with 98% concordance using a k-mer based method, can process 10,000 samples in just over 1 day using a standard server and can call serotypes at a coverage as low as 10x. SeroBA is implemented in Python3 and is freely available under an open source GPLv3 license from: https://github.com/sanger-pathogens/seroba\n\nDATA SUMMARYO_LIThe reference genome Streptococcus pneumoniae ATCC 700669 is available from National Center for Biotechnology Information (NCBI) with the accession number: FM211187\nC_LIO_LISimulated paired end reads for experiment 2 have been deposited in FigShare: https://doi.org/10.6084/m9.figshare.5086054.v1\nC_LIO_LIAccession numbers for all other experiments are listed in Supplementary Table S1 and Supplementary Table S2.\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTThis article describes SeroBA, a k-mer based method for predicting the serotypes of Streptococcus pneumoniae from Whole Genome Sequencing (WGS) data. SeroBA can identify 92 serotypes and 2 subtypes with constant memory usage and low computational costs. We showed that SeroBA is able to reliably predict serotypes at a depth of coverage as low as 10x and is scalable to large datasets.

bioinformatics

ARIBA: rapid antimicrobial resistance genotyping directly from sequencing reads

Antimicrobial resistance (AMR) is one of the major threats to human and animal health worldwide, yet few high-throughput tools exist to analyse and predict the resistance of a bacterial isolate from sequencing data. Here we present a new tool, ARIBA, that identifies AMR-associated genes and single nucleotide polymorphisms directly from short reads, and generates detailed and customisable output. The accuracy and advantages of ARIBA over other tools are demonstrated on three datasets from Gram-positive and Gram-negative bacteria, with ARIBA outperforming existing methods. ARIBA is available at https://github.com/sanger-pathogens/ariba.

bioinformatics

Comparison Of Multi-locus Sequence Typing Software For Next Generation Sequencing Data

Multi-locus sequence typing (MLST) is a widely used method for categorising bacteria. Increasingly MLST is being performed using next generation sequencing data by reference labs and for clinical diagnostics. Many software applications have been developed to calculate sequence types from NGS data; however, there has been no comprehensive review to date on these methods. We have compared six of these applications against real and simulated data and present results on: 1. the accuracy of each method against traditional typing methods, 2. the performance on real outbreak datasets, 3. in the impact of contamination and varying depth of coverage, and 4. the computational resource requirements.\n\nDATA SUMMARYO_LISimulated reads for datasets testing coverage and mixed samples have been deposited in Figshare; DOI: https://doi.org/10.6084/m9.figshare.4602301.vl\nC_LIO_LIOutbreak databases are available from Github; url - https://github.com/WGS-standards-and-analysis/datasets\nC_LIO_LIDocker containers used to run each of the applications are available from Github; url - https://tinyurl.com/z7ks2ft\nC_LIO_LIAccession numbers for the data used in this paper are available in the Supplementary material.\nC_LI\n\nWe confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {ballotx}\n\nIMPACT STATEMENTSequence typing is rapidly transitioning from traditional sequencing methods to using whole genome sequencing. A number of in silico prediction methods have been developed on an ad hoc basis and aim to replicate Multi-locus sequence typing (MLST). This is the first study to comprehensively evaluate multiple MLST software applications on real validated datasets and on common simulated difficult cases. It will give researchers a clearer understanding of the accuracy, limitations and computational performance of the methods they use, and will assist future researchers to choose the most appropriate method for their experimental goals.

bioinformatics