bioRxiv ScienceSearch

Biology subjects

Corander, J.

Publications and source records attributed to Corander, J..

17 recordsLinked to original sources

Fast Hierarchical Bayesian Analysis of Population Structure

We present fastbaps, a fast solution to the genetic clustering problem. Fastbaps rapidly identifies an approximate fit to a Dirichlet Process Mixture model (DPM) for clustering multilocus genotype data. Our efficient model-based clustering approach is able to cluster datasets 10-100 times larger than the existing model-based methods, which we demonstrate by analysing an alignment of over 110,000 sequences of HIV-1 pol genes. We also provide a method for rapidly partitioning an existing hierarchy in order to maximise the DPM model marginal likelihood, allowing us to split phylogenetic trees into clades and subclades using a population genomic model. Extensive tests on simulated data as well as a diverse set of real bacterial and viral datasets show that fastbaps provides comparable or improved solutions to previous model-based methods, while generally being significantly faster. The method is made freely available under an open source MIT licence as an easy to use R package at https://github.com/gtonkinhill/fastbaps.

genomics

Rapid statistical methods for inferring intra- and inter-hospital transmission of nosocomial pathogens from whole genome sequence data

Whole genome sequence (WGS) data for bacterial pathogens can provide evidence as to the source of nosocomial infection, and more specifically the ability to distinguish between intra- and inter-hospital transmission. This is currently achieved either through using SNP thresholds, which can lack statistical robustness, or by constructing phylogenetic trees, which can be computationally expensive and difficult to interpret. Here we compare two alternative statistical approaches using 1022 genomes of methicillin resistant Staphylococcus aureus (MRSA) clone ST22. In 71% of cases both methods predict the same hospital origin, which is also supported by the ML tree. Robust assignments are divided approximately equally between intra-hospital transmission and inter-hospital transmission. Our approaches are rapid and produce intuitive output that could inform on immediate infection control priorities, as well as providing long-term data on inter-hospital transmission networks. We discuss the strengths and weakness of our methods, and the generalisability of this approach.\n\nOne Sentence SummaryWe present rapid statistical methods for distinguishing intra- versus inter-hospital transmission of bacterial pathogens using whole genome sequence data; these methods do not require the use of SNP thresholds or the generation and interpretation of phylogenetic trees.

epidemiology

Prediction of post-vaccine population structure of Streptococcus pneumoniae using accessory gene frequencies

Predicting how pathogen populations will change over time is challenging. Such has been the case with Streptococcus pneumoniae, an important human pathogen, and the pneumococcal conjugate vaccines (PCVs), which target only a fraction of the strains in the population. Here, we use the frequencies of accessory genes to predict changes in the pneumococcal population after vaccination, hypothesizing that these frequencies reflect negative frequency-dependent selection (NFDS) on the gene products. We find that the standardized predicted fitness of a strain estimated by an NFDS-based model at the time the vaccine is introduced enables to predict whether the strain increases or decreases in prevalence following vaccination. Further, we are able to forecast the equilibrium post-vaccine population composition and assess the invasion capacity of emerging lineages. Overall, we provide a method for predicting the impact of an intervention on pneumococcal populations with potential application to other bacterial pathogens in which NFDS is a driving force.

evolutionary biology

Signatures of negative frequency dependent selection in colonisation factors and the evolution of a multi-drug resistant lineage of Escherichia coli

Escherichia coli is a major cause of bloodstream and urinary tract infections globally. The wide dissemination of multi-drug resistant (MDR) strains of extra-intestinal pathogenic E. coli (ExPEC) poses a rapidly increasing public health burden due to narrowed treatment options and increased risk of failure to clear an infection. Here, we present a detailed population genomic analysis of the ExPEC ST131 clone, in which we seek explanations for its success as an emerging pathogenic strain beyond the acquisition of antimicrobial resistance (AMR) genes. We show evidence for evolution towards separate ecological niches for the main clades of ST131 and differential evolution of anaerobic metabolism, key colonisation and virulence factors. We further demonstrate that negative frequency-dependent selection acting across accessory loci is a major mechanism that has shaped the population evolution of this pathogen.

microbiology

Fast and flexible bacterial genomic epidemiology with PopPUNK

The routine use of genomics for disease surveillance provides the opportunity for high-resolution bacterial epidemiology.\n\nHowever, current whole-genome clustering and multi-locus typing approaches do not fully exploit core and accessory genomic variation, and cannot both automatically identify, and subsequently expand, clusters of significantly-similar isolates in large datasets and across species.\n\nHere we describe PopPUNK (Population Partitioning Using Nucleotide K-mers; https://poppunk.readthedocs.io/en/latest/). software implementing scalable and expandable annotation- and alignment-free methods for population analysis and clustering.\n\nVariable-length k-mer comparisons are used to distinguish isolates divergence in shared sequence and gene content, which we demonstrate to be accurate over multiple orders of magnitude using both simulated data and real datasets from ten taxonomically-widespread species. Connections between closely-related isolates of the same strain are robustly identified, despite variation in the discontinuous pairwise distance distributions that reflects species diverse evolutionary patterns. PopPUNK can process 103-104 genomes as single batch, with minimal memory use and runtimes up to 200-fold faster than existing methods. Clusters of strains remain consistent as new batches of genomes are added, which is achieved without needing to re-analyse all genomes de novo.\n\nThis facilitates real-time surveillance with stable cluster naming and allows for outbreak detection using hundreds of genomes in minutes. Interactive visualisation and online publication is streamlined through automatic output of results to multiple platforms.\n\nPopPUNK has been designed as a flexible platform that addresses important issues with currently used whole-genome clustering and typing methods, and has potential uses across bacterial genetics and public health research.

genomics

High-resolution sweep metagenomics using ultrafast read mapping and inference

Determining the composition of bacterial communities beyond the level of a genus or species is challenging because of the considerable overlap between genomes representing close relatives. Here, we present the mSWEEP method for identifying and estimating the relative abundances of bacterial lineages from plate sweeps of enrichment cultures. mSWEEP leverages biologically grouped sequence assembly databases, applying probabilistic modelling, and provides controls for false positive results. Using sequencing data from major pathogens, we demonstrate significant improvements in lineage quantification and detection accuracy. Our method facilitates investigating cultures comprising mixtures of bacteria, and opens up a new field of plate sweep metagenomics.

bioinformatics

Antimicrobial exposure in sexual networks drives divergent evolution in modern gonococci

The sexually transmitted pathogen Neisseria gonorrhoeae is regarded as being on the way to becoming an untreatable superbug. Despite its clinical importance, little is known about its emergence and evolution, and how this corresponds with the introduction of antimicrobials. We present a genome-based phylogeographic analysis of 419 gonococcal isolates from across the globe. Results indicate that modern gonococci originated in Europe or Africa as late as the 16thcentury and subsequently disseminated globally. We provide evidence that the modern gonococcal population has been shaped by antimicrobial treatment of sexually transmitted and other infections, leading to the emergence of two major lineages with different evolutionary strategies. The well-described multi-resistant lineage is associated with high rates of homologous recombination and infection in high-risk sexual networks where antimicrobial treatment is frequent. A second, multi-susceptible lineage associated with heterosexual networks, where asymptomatic infection is more common, was also identified, with potential implications for infection control.

genomics

mlplasmids: a user-friendly tool to predict plasmid- and chromosome-derived sequences for single species

Assembly of bacterial short-read whole genome sequencing (WGS) data frequently results in hundreds of contigs for which the origin, plasmid or chromosome, is unclear. Long-read sequencing has emerged as a solution to resolve plasmid structures and to obtain complete genomes for most bacterial species. This information can be used to generate and label datasets from short-read based contigs as plasmid- or chromosome-derived. We investigated the use of several popular machine learning methods to classify short-read contigs with known plasmid- or chromosome-origin from Enterococcus faecium, Klebsiella pneumoniae and Escherichia coli using pentamer frequencies. Based on resulting F1-scores we selected support-vector machine (SVM) models as best classifier for all three bacterial species (F1-score E. faecium = 0.94, F1-score K. pneumoniae = 0.90, F1-score E. coli = 0.76), which outperformed other existing plasmid tools using an independent set of isolates (precision E. faecium = 0.92, precision K. pneumoniae = 0.86, precision E. coli = 0.82). We demonstrated the scalability of our model by accurately predicting the plasmidome of a large collection of 1,644 E. faecium isolates with only short-read WGS available using a standard laptop with a single core. A low number of false positive predicted sequences suggests that the assignment of a particular gene of interest as plasmid- or chromosome-encoded by the models is plausible. The SVM classifiers are publicly available as a new R package called mlplasmids at https://gitlab.com/sirarredondo/mlplasmids under the GNU General Public License v3.0. We additionally developed a graphical-user interface using the Shiny package which can be accessed at https://sarredondo.shinyapps.io/mlplasmids/. Single genomes can easily be predicted by uploading genome assemblies. We anticipate that this tool may significantly facilitate research on the dissemination of plasmids encoding antibiotic resistance and/or contributing to host adaptation.

microbiology

Genomic determinants of sympatric speciation of the Mycobacterium tuberculosis complex across evolutionary timescales.

BACKGROUNDModels on how bacterial lineages differentiate increase our understanding on early bacterial speciation events and about the genetic loci involved. Here, we analyze the population genomics events leading to the emergence of the tuberculosis pathogen.\n\nRESULTSThe emergence is characterized by a combination of recombination events involving core pathogenesis functions and purifying selection on early diverging loci. We identify the phoR gene, the sensor kinase of a two-component system involved in virulence, as a key functional player subject to pervasive positive selection after the divergence of the MTBC from its ancestor. Previous evidence showed that phoR mutations played a central role in the adaptation of the pathogen to different host species. Now we show that phoR have been under selection during the early spread of human tuberculosis, during later expansions and in on-going transmission events.\n\nCONCLUSIONSOur results show that linking pathogen evolution across evolutionary and epidemiological timescales point to past and present virulence determinants.

evolutionary biology

pyseer: a comprehensive tool for microbial pangenome-wide association studies

SummaryGenome-wide association studies (GWAS) in microbes face different challenges to eukaryotes and have been addressed by a number of different methods. pyseer brings these techniques together in one package tailored to microbial GWAS, allows greater flexibility of the input data used, and adds new methods to interpret the association results.\n\nAvailability and Implementationpyseer is written in python and is freely available at https://github.com/mgalardini/pyseer, or can be installed through pip. Documentation and a tutorial are available at http://pyseer.readthedocs.io.\n\nContactjohn.lees@nyumc.org and marco@ebi.ac.uk\n\nSupplementary informationSupplementary data are available online.

bioinformatics

Resolving outbreak dynamics using Approximate Bayesian Computation for stochastic birth-death models

Earlier research has suggested that Approximate Bayesian Computation (ABC) makes it possible to fit simulator-based intractable birth-death models to investigate communicable disease outbreak dynamics with accuracy comparable to that of exact Bayesian methods. However, recent findings have indicated that key parameters such as the reproductive number R may remain poorly identifiable. Here we show that the identifiability issue can be resolved by taking into account disease-specific characteristics of the transmission process in closer detail. Using tuberculosis (TB) in the San Francisco Bay area as a case-study, we consider the situation where the genotype data are generated as a mixture of three stochastic processes, each with their distinct dynamics and clear epidemiological interpretation.\n\nThe ABC inference yields stable and accurate posterior inferences about outbreak dynamics from aggregated annual case data with genotype information. We also show that under the proposed model, the infectious population size can be reliably inferred from the data. The estimate is approximately two orders of magnitude smaller compared to assumptions made in the earlier ABC studies, and is much better aligned with epidemiological knowledge about active TB prevalence. Similarly, the reproductive number R related to the primary underlying transmission process is estimated to be nearly three-fold compared with the previous estimates, which has a substantial impact on the interpretation of the fitted outbreak model.

bioinformatics

SuperDCA for genome-wide epistasis analysis

The potential for genome-wide modeling of epistasis has recently surfaced given the possibility of sequencing densely sampled populations and the emerging families of statistical interaction models. Direct coupling analysis (DCA) has earlier been shown to yield valuable predictions for single protein structures, and has recently been extended to genome-wide analysis of bacteria, identifying novel interactions in the co-evolution between resistance, virulence and core genome elements. However, earlier computational DCA methods have not been scalable to enable model fitting simultaneously to 104-105 polymorphisms, representing the amount of core genomic variation observed in analyses of many bacterial species. Here we introduce a novel inference method (SuperDCA) which employs a new scoring principle, efficient parallelization, optimization and filtering on phylogenetic information to achieve scalability for up to 105 polymorphisms. Using two large population samples of Streptococcus pneumoniae, we demonstrate the ability of SuperDCA to make additional significant biological findings about this major human pathogen. We also show that our method can uncover signals of selection that are not detectable by genome-wide association analysis, even though our analysis does not require phenotypic measurements. SuperDCA thus holds considerable potential in building understanding about numerous organisms at a systems biological level.\n\nAuthor SummaryRecent work has demonstrated the emerging potential in statistical genome-wide modeling to uncover co-selection and epistatic interactions between polymorphisms in bacterial chromosomes from densely sampled population data. Here we develop the Potts model based approach further into a fully mature computational method which can be applied to most existing bacterial population genomic data sets in a straightforward manner. Our advances are relying on more efficient parameter scoring, highly optimized and parallelized open source C++ code, which does not rely on the computation-intensive polymorphism subsampling approximations used earlier. By analyzing the two largest available population samples of Streptococcus pneumoniae (the pneumococcus), we highlight several biological discoveries related to the survival of the pneumococcus and co-evolution of penicillin-binding loci, which were not uncovered by the earlier analyses. Our method holds considerable potential for building understanding about numerous organisms at a systems biological level.

genomics

PANINI: Pangenome Neighbor Identification for Bacterial Populations

The standard workhorse for genomic analysis of the evolution of bacterial populations is phylogenetic modelling of mutations in the core genome. However, in the current era of population genomics, a notable amount of information about evolutionary and transmission processes in diverse populations can be lost unless the accessory genome is also taken into consideration. Here we introduce PANINI, a computationally scalable method for identifying the neighbours for each isolate in a data set using unsupervised machine learning with stochastic neighbour embedding. PANINI is browser-based and integrates with the Microreact platform for rapid online visualisation and exploration of both core and accessory genome evolutionary signals together with relevant epidemiological, geographic, temporal and other metadata. Several case studies with single-and multi-clone pneumococcal populations are presented to demonstrate ability to identify biologically important signals from gene content data. PANINI is available at http://panini.wgsa.net/ and code at http://gitlab.com/cgps/panini

microbiology

Bacmeta: simulation for genomic evolution in bacterial metapopulations

The advent of genomic data from densely sampled bacterial populations has created a need for flexible simulators by which models and hypotheses can be efficiently investigated in the light of empirical observations. Bacmeta provides fast stochastic simulation of neutral evolution within a large collection of interconnected bacterial populations with completely adjustable connectivity network. Stochastic events of mutations, recombinations, insertions/deletions, migrations and microepidemics can be simulated in discrete non-overlapping generations with a Wright-Fisher model that operates on explicit sequence data of any desired genome length. Each model component, including locus, bacterial strain, population, and ultimately the whole metapopulation, is efficiently simulated using C++ objects, and detailed metadata from each level of the simulation can be acquired. The software can be executed in a cluster environment using simple textual input files, enabling, e.g., large-scale simulations and likelihood-free inference. Bacmeta is implemented with C++ for Linux, Mac and Windows. It is available at https://bitbucket.org/aleksisipola/bacmeta under the BSD 3-clause license.\n\nContactaleksi.sipola@helsinki.fi,\n\njukka.corander@medisin.uio.no\n\nSupplementary informationSupplementary data are available online at bioRxiv.

genomics

The contribution of genetic variation of Streptococcus pneumoniae to the clinical manifestation of invasive pneumococcal disease

BackgroundDifferent clinical manifestations of invasive pneumococcal disease (IPD) have thus far mainly been explained by patient characteristics. Here we studied the contribution of pneumococcal genetic variation to IPD phenotype.\n\nMethodsThe index cohort consisted of 349 patients admitted to two Dutch hospitals between 2000-2011 with pneumococcal bacteraemia. We performed genome-wide association studies to identify pneumococcal lineages, genes and allelic variants associated with 23 clinical IPD phenotypes. The identified associations were validated in a nationwide (n=482) and a post-pneumococcal vaccination cohort (n=121). The contribution of confirmed pneumococcal genotypes to the clinical IPD phenotype, relative to known clinical predictors, was tested by regression analysis.\n\nFindingsThe presence of pneumococcal gene slaA was a nationwide confirmed independent predictor of meningitis (OR=10.5, p=0.001), as was sequence cluster 9 (OR=3.68, p=0.057). A set of 4 pneumococcal genes co-located on a prophage was a confirmed independent predictor of 30-day mortality (OR=3.4, p=0.003). We could detect the pneumococcal variants of concern in these patients blood samples by molecular amplification. In the post-vaccination cohort where the distribution of both patient characteristics and pneumococcal serotypes had changed, the relative importance of the prophage was no longer supported.\n\nInterpretationKnowledge of pneumococcal genotypic variants improved our clinical risk assessment for detrimental manifestations of IPD. This provides us with novel opportunities to target, anticipate or avert the pathogenic effects that are related to particular pneumococcal variants. Therefore, future diagnostics should facilitate prompt appreciation of pathogen diversity in clinical sepsis management. Ongoing surveillance is warranted to monitor the clinical value of information on pathogen variants in dynamic microbial and susceptible host populations.\n\nFundingNone.

microbiology

Phylogeographic separation and formation of sexually discreet lineages in a global population of Yersinia pseudotuberculosis

Yersinia pseudotuberculosis is a Gram negative intestinal pathogen of humans and has been responsible for several nation-wide gastro-intestinal outbreaks. Large-scale population genomic studies have been performed on the other human pathogenic Yersinia, Y. pestis and Y. enterocolitica allowing a high-resolution understanding of the ecology, evolution and dissemination of these pathogens. However, to date no large-scale global population genomic analysis of Y. pseudotuberculosis has been performed. Here we present analyses of the genomes of 134 strains of Y. pseudotuberculosis isolated from around the world, from multiple ecosystems since 1960s. Our data display a phylogeographic split within the population, with an Asian ancestry and subsequent dispersal of successful clonal lineages into Europe and the rest of the world. These lineages can be differentiated by CRISPR cluster arrays, and we show that the lineages are limited with respect to inter-lineage genetic exchange. This restriction of genetic exchange maintains the discrete lineage structure in the population despite co-existence of lineages for thousands of years in multiple countries. Our data highlights how CRISPR can be informative of the evolutionary trajectory of bacterial lineages, and merits further study across bacteria.

microbiology

Weak epistasis may drive adaptation in recombining bacteria

The impact of epistasis on the evolution of multilocus traits depends on recombination. Population genetic theory has been largely developed for eukaryotes, many of which recombine so frequently that epistasis between polymorphisms has not been considered to play a large role in adaptation and has been compared to the fleeting influence of non-heritable effects. Many bacteria also recombine, some to the degree that their populations are described as panmictic or freely recombining. However, whether this recombination is sufficient to limit the ability of selection to act on epistatic contributions to fitness is unknown. We create a sensitive method to quantify homologous recombination in five bacterial pathogens and use these parameter estimates in a multilocus model of bacterial evolution with additive and epistatic effects. We find that even for highly recombining species (e.g. Streptococcus pneumoniae or Helicobacter pylori), selection may act on the cumulative effects of weak (as well as strong) interactions between distant mutations since homologous recombination typically transfers only short segments. Furthermore, whether selection acts more efficiently on physically proximal loci depends on the average recombination tract length. Epistasis may thus play an important role in the adaptive evolution of bacteria and, unlike in eukaryotes, does not need to be strong, involve near loci, or require specific metapopulation dynamics.

evolutionary biology