bioRxiv ScienceSearch

Biology subjects

Julian Parkhill

Publications and source records attributed to Julian Parkhill.

10 recordsLinked to original sources

Interacting networks of resistance, virulence and core machinery genes identified by genome-wide epistasis analysis

Recent advances in the scale and diversity of population genomic datasets for bacteria now provide the potential for genome-wide patterns of co-evolution to be studied at the resolution of individual bases. The major human pathogen Streptococcus pneumoniae represents the first bacterial organism for which densely enough sampled population data became available for such an analysis. Here we describe a new statistical method, genomeDCA, which uses recent advances in computational structural biology to identify the polymorphic loci under the strongest co-evolutionary pressures. Genome data from over three thousand pneumococcal isolates identified 5,199 putative epistatic interactions between 1,936 sites. Over three-quarters of the links were between sites within the pbp2x, pbp1a and pbp2b genes, the sequences of which are critical in determining non-susceptibility to beta-lactam antibiotics. A network-based analysis found these genes were also coupled to that encoding dihydrofolate reductase, changes to which underlie trimethoprim resistance. Distinct from these resistance genes, a large network component of 384 protein coding sequences encompassed many genes critical in basic cellular functions, while another distinct component included genes associated with virulence. These results have the potential both to identify previously unsuspected protein-protein interactions, as well as genes making independent contributions to the same phenotype. This approach greatly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.\n\nAuthor SummaryEpistatic interactions between polymorphisms in DNA are recognized as important drivers of evolution in numerous organisms. Study of epistasis in bacteria has been hampered by the lack of both densely sampled population genomic data, suitable statistical models and powerful inference algorithms for extremely high-dimensional parameter spaces. We introduce the first model-based method for genome-wide epistasis analysis and use the largest available bacterial population genome data set on Streptococcus pneumoniae (the pneumococcus) to demonstrate its potential for biological discovery. Our approach reveals interacting networks of resistance, virulence and core machinery genes in the pneumococcus, which highlights putative candidates for novel drug targets. Our method significantly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.

Genetics

Large scale genomic analysis shows no evidence for repeated pathogen adaptation during the invasive phase of bacterial meningitis in humans

Recent studies have provided evidence for rapid pathogen genome variation, some of which could potentially affect the course of disease. We have previously detected such variation by comparing isolates infecting the blood and cerebrospinal fluid (CSF) of a single patient during a case of bacterial meningitis.\n\nTo determine whether the observed variation repeatedly occurs in cases of disease, we performed whole genome sequencing of paired isolates from blood and CSF of 938 meningitis patients. We also applied the same techniques to 54 paired isolates from the nasopharynx and CSF.\n\nUsing a combination of reference-free variant calling approaches we show that no genetic adaptation occurs in the invasive phase of bacterial meningitis for four major pathogen species: Streptococcus pneumoniae, Neisseria meningitidis, Listeria monocytogenes and Haemophilus influenzae. From nasopharynx to CSF, no adaptation was seen in S. pneumoniae, but in N. meningitidis mutations potentially mediating adaptation to the invasive niche were occasionally observed in the dca gene.\n\nThis study therefore shows that the bacteria capable of causing meningitis are already able to do this upon entering the blood, and no further sequence change is necessary to cross the blood-brain barrier. The variation discovered from nasopharyngeal isolates suggest that larger studies comparing carriage and invasion may help determine the likely mechanisms of invasiveness.\n\nAuthor SummaryWe have analysed the entire DNA sequence from bacterial pathogen isolates from cases of meningitis in 938 Dutch adults, focusing on comparing pairs of isolates from the patients blood and their cerebrospinal fluid. Previous research has been on only a single patient, but showed possible signs of adaptation to treatment within the host over the course of a single case of disease.\n\nBy sequencing many more such paired samples, and including four different bacterial species, we were able to determine that adaptation of the pathogen does not occur after bloodstream invasion during bacterial meningitis.\n\nWe also analysed 54 pairs of isolates from pre- and post-invasive niches from the same patient. In N. meningitidis we found variation in the sequence of one gene which appears to provide bacteria with an advantage after invasion of the bloodstream.\n\nOverall, our findings indicate that evolution after invasion in bacterial meningitis is not a major contribution to disease pathogenesis. Future studies should involve more extensive sampling between the carriage and disease niches, or on variation of the host.

Genomics

Genomic dissection of an Icelandic epidemic of equine respiratory disease

The native horse population of Iceland has remained free of major infectious diseases. Between May and July 2010 an epidemic of respiratory disease swept through the population. Initial microbiological investigations ruled out known equine viral agents as the cause of the infections, but identified the opportunistic pathogen Streptococcus zooepidemicus as being frequently isolated from diseased animals. This diverse bacterial species has a broad host range and is usually regarded as a commensal of horses. By genome sequencing S. zooepidemicus recovered from horses during the epidemic we show that although multiple clones of S. zooepidemicus were present in the population, one particular clone, ST209, was responsible for the epidemic. Concurrent with the epidemic, ST209 caused zoonotic infections, highlighting the pathogenic potential of this clone. Phylogenetic analysis suggests that the original ST209 strain entered Iceland in late 2008 or early 2009. Epidemiological investigation revealed that the incursion of this strain into a training yard that utilized a submerged treadmill between the 5th and 19th of February 2010 was a critical trigger for the ensuing epidemic of disease, provided a nidus for the infection of multiple horses, and subsequent distribution of these animals to multiple sites in Iceland.

Microbiology

Comparison of bacterial genome assembly software for MinION data

Antimicrobial resistance genes can be carried on plasmids or on mobile elements integrated into the chromosome. We sequenced a multidrug resistant Enterobacter kobei genome isolated from wastewater in the United Kingdom, but were unable to conclusively identify plasmids from the short read assembly. Our aim was to compare and contrast the accuracy and characteristics of open source software (PBcR, Canu, miniasm and SPAdes) for the assembly of bacterial genomes (including plasmids) generated by the MinION instrument. Miniasm produced an assembly in the shortest time, but Canu produced the most accurate assembly overall. We found that MinION data alone was able to generate a contiguous and accurate assembly of an isolate with multiple plasmids.

Genomics

Robust high throughput prokaryote de novo assembly and improvement pipeline for Illumina data

The rapidly reducing cost of bacterial genome sequencing has lead to its routine use in large scale microbial analysis. Though mapping approaches can be used to find differences relative to the reference, many bacteria are subject to constant evolutionary pressures resulting in events such as the loss and gain of mobile genetic elements, horizontal gene transfer through recombination and genomic rearrangements. De novo assembly is the reconstruction of the underlying genome sequence, an essential step to understanding bacterial genome diversity. Here we present a high throughput bacterial assembly and improvement pipeline that has been used to generate nearly 20,000 draft genome assemblies in public databases. We demonstrate its performance on a public data set of 9,404 genomes. We find all the genes used in MLST schema present in 99.6% of assembled genomes. When tested on low, neutral and high GC organisms, more than 94% of genes were present and completely intact. The pipeline has proven to be scalable and robust with a wide variety of datasets without requiring human intervention. All of the software is available on GitHub under the GNU GPL open source license.\n\nDATA SUMMARYO_LIThe assembly pipeline software is available from Github under the GNU GPL open source license; (url - https://github.com/sanger-pathogens/vr-codebase)\nC_LIO_LIThe assembly improvement software is available from Github under the GNU GPL open source license; (url - https://github.com/sanger-pathogens/assembly_improvement)\nC_LIO_LIAccession numbers for 9,404 assemblies are provided in the supplementary material.\nC_LIO_LIThe Bordetella pertussis sample has sample accession ERS1058649, sequencing reads accession ERR1274624 and assembly accessions FJMX01000001-FJMX01000249.\nC_LIO_LIThe Salmonella enterica subsp. enterica serovar Pullorum sample has sample accession ERS1058652, sequencing reads accession ERR1274625 and assembly accession FJMV01000001-FJMV01000026.\nC_LIO_LIThe Staphylococcus aureus sample has sample accession ERS1058648, sequencing reads accession ERR1274626 and assembly accessions FJMW01000001-FJMW01000040.\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files.{ballotcheck}\n\nIMPACT STATEMENTThe pipeline described in this paper has been used to assemble and annotate 30% of all bacterial genome assemblies in GenBank (18,080 out of 59,536, accessed 16/2/16). The automated generation of de novo assemblies is a critical step to explore bacterial genome diversity. MLST genes are found in 99.6% of cases, making it at least as good as existing typing methods. In the test genomes we present, more than 94% of genes are correctly assembled into intact reading frames.

Bioinformatics

Sequence element enrichment analysis to determine the genetic basis of bacterial phenotypes

Bacterial genomes vary extensively in terms of both gene content and gene sequence - this plasticity hampers the use of traditional SNP-based methods for identifying all genetic associations with phenotypic variation. Here we introduce a computationally scalable and widely applicable statistical method (SEER) for the identification of sequence elements that are significantly enriched in a phenotype of interest. SEER is applicable to even tens of thousands of genomes by counting variable-length k-mers using a distributed string-mining algorithm. Robust options are provided for association analysis that also correct for the clonal population structure of bacteria. Using large collections of genomes of the major human pathogens Streptococcus pneumoniae and Streptococcus pyogenes, SEER identifies relevant previously characterised resistance determinants for several antibiotics and discovers potential novel factors related to the invasiveness of S. pyogenes. We thus demonstrate that our method can answer important biologically and medically relevant questions.

Genomics

Circlator: automated circularization of genome assemblies using long sequencing reads

The assembly of DNA sequence data into finished genomes is undergoing a renais-sance thanks to emerging technologies producing reads of tens of kilobases. Assembling complete bacterial and small eukaryotic genomes is now possible, but the final step of circularizing sequences remains unsolved. Here we present Circlator, the first tool to automate assembly circularization and produce accurate linear rep-resentations of circular sequences. Using Pacific Biosciences and Oxford Nanopore data, Circlator correctly circularized 26 of 27 circularizable sequences, comprising 11 chromosomes and 12 plasmids from bacteria, the apicoplast and mitochondrion of Plasmodium falciparum and a human mitochondrion. Circlator is available at http://sanger-pathogens.github.io/circlator/.

Bioinformatics

Roary: Rapid large-scale prokaryote pan genome analysis

SummaryA typical prokaryote population sequencing study can now consist of hundreds or thousands of isolates. Interrogating these datasets can provide detailed insights into the genetic structure of of prokaryotic genomes. We introduce Roary, a tool that rapidly builds large-scale pan genomes, identifying the core and dispensable accessory genes. Roary makes construction of the pan genome of thousands of prokaryote samples possible on a standard desktop without compromising on the accuracy of results. Using a single CPU Roary can produce a pan genome consisting of 1000 isolates in 4.5 hours using 13 GB of RAM, with further speedups possible using multiple processors.\n\nAvailability and implementationRoary is implemented in Perl and is freely available under an open source GPLv3 license from http://sanger-pathogens.github.io/Roary\n\nContactroary@sanger.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Genome specialization and decay of the strangles pathogen, Streptococcus equi, is driven by persistent infection

Strangles, the most frequently diagnosed infectious disease of horses worldwide, is caused by Streptococcus equi. Despite its prevalence, the global diversity and mechanisms underlying the evolution of S. equi as a host-restricted pathogen remain poorly understood. Here we define the global population structure of this important pathogen and reveal a population replacement in the late 19th or early 20th century, contemporaneous with a spate of global conflicts. Our data reveal a dynamic genome that continues to mutate and decay, but also to amplify and acquire genes despite the organism having lost its natural competence and become host-restricted.\n\nThe lifestyle of S. equi within the horse is defined by short-term acute disease, strangles, followed by long-term carriage. Population analysis reveals evidence of convergent evolution in isolates from post-acute disease samples, as a result of niche adaptation to persistent carriage within a host. Mutations that lead to metabolic streamlining and the loss of virulence determinants are more frequently found in carriage isolates, suggesting that the pathogenic potential of S. equi reduces as a consequence of long term residency within the horse post acute disease. An example of this is the deletion of the equibactin siderophore locus that is associated with iron acquisition, which occurs exclusively in carrier isolates, and renders S. equi significantly less able to cause acute disease in the natural host. We identify several loci that may similarly be required for the full virulence of S. equi, directing future research towards the development of new vaccines against this host-restricted pathogen.

Microbiology

Reagent contamination can critically impact sequence-based microbiome analyses

The study of microbial communities has been revolutionised in recent years by the widespread adoption of culture independent analytical techniques such as 16S rRNA gene sequencing and metagenomics. One potential confounder of these sequence-based approaches is the presence of contamination in DNA extraction kits and other laboratory reagents. In this study we demonstrate that contaminating DNA is ubiquitous in commonly used DNA extraction kits, varies greatly in composition between different kits and kit batches, and that this contamination critically impacts results obtained from samples containing a low microbial biomass. Contamination impacts both PCR based 16S rRNA gene surveys and shotgun metagenomics. These results suggest that caution should be advised when applying sequence-based techniques to the study of microbiota present in low biomass environments. We provide an extensive list of potential contaminating genera, and guidelines on how to mitigate the effects of contamination. Concurrent sequencing of negative control samples is strongly advised.

Molecular Biology