bioRxiv ScienceSearch

Biology subjects

Paul Keim

Publications and source records attributed to Paul Keim.

6 recordsLinked to original sources

A Bacillus anthracis Genome Sequence from the Sverdlovsk 1979 Autopsy Specimens

Anthrax is a zoonotic disease that occurs naturally in wild and domestic animals but has been used by both state-sponsored programs and terrorists as a biological weapon. The 2001 anthrax letter attacks involved less than gram quantities of Bacillus anthracis spores while the earlier Soviet weapons program produced tons. A Soviet industrial production facility in Sverdlovsk proved deficient in 1979 when a plume of spores was accidentally released and resulted in one of the largest known human anthrax outbreak. In order to understand this outbreak and others, we have generated a B. anthracis population genetic database based upon whole genome analysis to identify all SNPs across a reference genome. Only ~12,000 SNPs were identified in this low diversity species and represents the breadth of its known global diversity. Phylogenetic analysis has defined three major clades (A, B and C) with B and C being relatively rare compared to A. The A clade has numerous subclades including a major polytomy named the Trans-Eurasian (TEA) group. The TEA radiation is a dominant evolutionary feature of B. anthracis, many contemporary populations, and must have resulted from large-scale dispersal of spores from a single source. Two autopsy specimens from the Sverdlovsk outbreak were deeply sequenced to produce draft B. anthracis genomes. This allowed the phylogenetic placement of the Sverdlovsk strain into a clade with two Asian live vaccine strains, including the Russian Tsiankovskii strain. The genome was examined for evidence of drug resistance manipulation or other genetic engineering, but none was found. Only 13 SNPs differentiated the virulent Sverdlovsk strain from its common ancestor with two vaccine strains. The Soviet Sverdlovsk strain genome is consistent with a wild type strain from Russia that had no evidence of genetic manipulation during its industrial production. This work provides insights into the world's largest biological weapons program and provides an extensive B. anthracis phylogenetic reference valuable for future anthrax investigations.\n\nImportanceThe 1979 Russian anthrax outbreak resulted from an industrial accident at the Soviet anthrax spore production facility in the city of Sverdlovsk. Deep genomic sequencing of two autopsy specimens generated a draft genome and phylogenetic placement of the Soviet Sverdlovsk anthrax strain. While it is known that Soviet scientists had genetically manipulated Bacillus anthracis, with the potential to evade vaccine prophylaxis and antibiotic therapeutics, there was no genomic evidence of this from the Sverdlovsk production strain genome. The whole genome SNP genotype of the Sverdlovsk strain was used to precisely identify it and its close relatives in the context of an extensive global B. anthracis strain collection. This genomic identity can now be used for forensic tracking of this weapons material on a global scale and for future anthrax investigations.

Genomics

Whole genome SNP typing to investigate methicillin-resistant Staphylococcus aureus carriage in a health-care provider as the source of multiple surgical site infections.

BackgroundPrevention of nosocomial transmission of infections is a central responsibility in the healthcare environment, and accurate identification of transmission events presents the first challenge. Phylogenetic analysis based on whole genome sequencing provides a high-resolution approach for accurately relating isolates to one another, allowing precise identification or exclusion of transmission events and sources for nearly all cases. We sequenced 24 methicillin-resistant Staphylococcus aureus (MRSA) genomes to retrospectively investigate a suspected point source of three surgical site infections (SSIs) that occurred over a one-year period. The source of transmission was believed to be a surgical team member colonized with MRSA, involved in all surgeries preceding the SSI cases, who was subsequently decolonized. Genetic relatedness among isolates was determined using whole genome single nucleotide polymorphism (SNP) data.\n\nResultsWhole genome SNP typing (WGST) revealed 283 informative SNPs between the surgical team members isolate and the closest SSI isolate. The second isolate was 286 and the third was thousands of SNPs different, indicating the nasal carriage strain from the surgical team member was not the source of the SSIs. Given the mutation rates estimated for S. aureus, none of the SSI isolates share a common ancestor within the past 14 years, further discounting any common point source for these infections. The decolonization procedures and resources spent on the point source infection control could have been prevented if WGST was performed at the time of the suspected transmission, instead of retrospectively.\n\nConclusionsWhole genome sequence analysis is an ideal method to exclude isolates involved in transmission events and nosocomial outbreaks, and coupling this method with epidemiological data can determine if a transmission event occurred. These methods promise to direct infection control resources more appropriately.

Genomics

KlebSeq: A Diagnostic Tool for Healthcare Surveillance and Antimicrobial Resistance Monitoring of Klebsiella pneumoniae

Healthcare-acquired infections (HAIs) kill tens of thousands of people each year and add significantly to healthcare costs. Multidrug resistant and epidemic strains are a large proportion of HAI agents, and multidrug resistant strains of Klebsiella pneumoniae, a leading HAI agent, have become an urgent public health crisis. In the healthcare environment, patient colonization of K. pneumoniae precedes infection, and transmission via colonization leads to outbreaks. Periodic patient screening for K. pneumoniae colonization has cost-effective and life-saving potential. In this study, we describe the design and validation of KlebSeq, a highly informative screening tool that detects Klebsiella species and identifies clinically important strains and characteristics using highly multiplexed amplicon sequencing without a live culturing step. We demonstrate the utility of this tool on several complex specimen types including urine, wound swabs and tissue, several types of respiratory, and fecal, showing K. pneumoniae species and clonal group identification and antimicrobial resistance and virulence profiling, including capsule typing. Use of this amplicon sequencing tool can be used to screen patients for K. pneumoniae carriage to assess risk of infection and outbreak potential, and the expansion of this tool can be used for several other HAI agents or applications.

Molecular Biology

The Northern Arizona SNP Pipeline (NASP): accurate, flexible, and rapid identification of SNPs in WGS datasets

Whole genome sequencing (WGS) of bacteria is becoming standard practice in many laboratories. Applications for WGS analysis include phylogeography and molecular epidemiology, using single nucleotide polymorphisms (SNPs) as the unit of evolution. The Northern Arizona SNP Pipeline (NASP) was developed as a reproducible pipeline that scales well with the large amount of WGS data typically used in comparative genomics applications. In this study, we demonstrate how NASP compares to other tools in the analysis of two real bacterial genomics datasets and one simulated dataset. Our results demonstrate that NASP produces comparable, and often better, results to other pipelines, but is much more flexible in terms of data input types, job management systems, diversity of supported tools, and output formats. We also demonstrate differences in results based on the choice of the reference genome and choice of inferring phylogenies from concatenated SNPs or alignments including monomorphic positions. NASP represents a source-available, version-controlled, unit-tested method and can be obtained from tgennorth.github.io/NASP.

Bioinformatics

Eighteenth century Yersinia pestis genomes reveal the long-term persistence of an historical plague focus

The 14th-18th century pandemic of Yersinia pestis caused devastating disease outbreaks in Europe for almost 400 years. The reasons for plagues persistence and abrupt disappearance in Europe are poorly understood, but could have been due to either the presence of now-extinct plague foci in Europe itself, or successive disease introductions from other locations. Here we present five Y. pestis genomes from one of the last European outbreaks of plague, from 1722 in Marseille, France. The lineage identified has not been found in any extant Y. pestis foci sampled to date, and has its ancestry in strains obtained from victims of the 14th century Black Death. These data suggest the existence of a previously uncharacterized historical plague focus that persisted for at least three centuries. We propose that this disease source may have been responsible for the many resurgences of plague in Europe following the Black Death.

Microbiology

Local population structure and patterns of Western Hemisphere dispersal for Coccidioides spp., the fungal cause of Valley Fever

Coccidioidomycosis (or Valley Fever) is a fungal disease with high morbidity and mortality that affects tens of thousands of people each year. This infection is caused by two sibling species, Coccidioides immitis and C. posadasii, which are endemic to specific arid locales throughout the Western Hemisphere, particularly the desert southwest of the United States. Recent epidemiological and population genetic data suggest that the geographic range of coccidioidomycosis is expanding as new endemic clusters have been identified in the state of Washington, well outside of the established endemic range. The genetic mechanisms and epidemiological consequences of this expansion are unknown and require better understanding of the population structure and evolutionary history of these pathogens. Here we perform multiple phylogenetic inference and population genomics analyses of 68 new and 18 previously published genomes. The results provide evidence of substantial population structure in C. posadasii and demonstrate presence of distinct geographic clades in central and southern Arizona as well as dispersed populations in Texas, Mexico, South America and Central America. Although a smaller number of C. immitis strains were included in the analyses, some evidence of phylogeographic structure was also detected in this species, which has been historically limited to California and Baja Mexico. Bayesian analyses indicated that C. posadasii is the more ancient of the two species and that Arizona contains the most diverse subpopulations. We propose a southern Arizona-northern Mexico origin for C. posadasii and describe a pathway for dispersal and distribution out of this region.

Genomics