bioRxiv Science⌕ Search

Biology subjects

Toussaint, J.

Publications and source records attributed to Toussaint, J..

3 recordsLinked to original sources

Rapid gene exchange explains differences in bacterial pangenome structure

The size and diversity of bacterial gene repertoires, known as pangenomes, vary widely across species. The evolutionary forces driving the maintenance of pangenomes is an open topic of debate, with contradictory theories suggesting that pangenomes exist as a result of neutral evolution, with all genes gained and lost at random, or that all genes provide a fitness benefit to the host and are maintained by positive selection. Modelling of pangenome dynamics has provided insight into how gene exchange explains observed gene frequency distributions, and stands as the only means of jointly inferring contributions of individual gene selection effects and mobility on the maintenance of pangenomes. However, previous modelling studies have not included both gene-level selection and mobility, and do not consider broadly sampled genome datasets for many species. To differentiate neutral and selective forces maintaining pangenomes, we developed a mechanistic model of gene-level evolution, Pansim, and a scalable model fitting framework, PopPUNK-mod. Together, these tools leverage rapid genome distance calculation to fit models of pangenome dynamics to datasets containing hundreds of thousands of genomes. We used this framework to compare the pangenome dynamics of over 400 different bacterial species, using over 600,000 genomes. We find that diversity in pangenome characteristics between species is driven predominantly by variation in the number of rapidly exchanged genes, while the rate of exchange of remaining genes is conserved. We find that bacterial phylogeny, rather than ecology, correlates with pangenome dynamics. We express that pan-species gene-level analyses are now needed to understand selection across accessory genes. Our work highlights the importance of gene exchange rate differences in governing differences in pangenome characteristics between species.

bioinformatics↗

BaGPipe: an automated, reproducible, and flexible pipeline for bacterial genome-wide association studies

Microbial genome-wide association study (GWAS) tools often require manual data processing steps, lack comprehensive workflows, and are limited by scalability issues, thus hindering the exploration of bacterial genetic traits. To address these challenges, we developed BaGPipe, an automated and flexible bacterial GWAS pipeline built using Nextflow and incorporating Pyseer for association analysis. BaGPipe integrates all essential components of a bacterial GWAS--spanning pre-processing, statistical analysis, and downstream visualisation--into a unified workflow that is reproducible and easy to deploy across diverse computational environments. BaGPipe was validated on a publicly available dataset of Streptococcus pneumoniae whole-genome sequences, and reproduced published findings with improved computational efficiency. BaGPipe was then applied to a dataset of Staphylococcus aureus whole-genome sequences, successfully identifying known and novel antibiotic resistance associations. By offering an accessible, efficient, and reproducible platform, BaGPipe accelerates bacterial GWAS and facilitates deeper exploration into the genetic underpinnings of phenotypic traits. Impact StatementThe increasing availability of bacterial genome sequences has created an opportunity for robust, reproducible tools to facilitate the discovery of novel genotype-phenotype associations. Despite the demonstrated utility of genome-wide association studies (GWAS) in identifying genetic determinants of disease, toxicity and antibiotic resistance, existing tools for bacterial GWAS often involve fragmented workflows requiring extensive manual intervention, limiting their adoption and reproducibility. Here, we introduce BaGPipe, a fully integrated bacterial GWAS pipeline that automates pre-processing, statistical analysis, and visualisation, thereby streamlining the entire workflow. With its flexibility, scalability, and ease of use, BaGPipe makes bacterial GWAS more accessible to researchers, enabling faster and more reliable insights into microbial genetics. This is an important step towards overcoming the computational and logistical barriers that have constrained bacterial GWAS, ultimately accelerating research into microbial evolution, resistance mechanisms, and the genetic basis of other key phenotypic traits. Data SummaryBaGPipe is freely available at https://github.com/sanger-pathogens/BaGPipe. The Streptococcus pneumoniae input dataset is available from the Pyseer tutorial (https://pyseer.readthedocs.io/en/master/tutorial.html#). The Staphylococcus aureus sequencing assemblies can be sourced from their ERS accession numbers provided in supplementary data. The reference assemblies, listed in the supplementary, can be sourced from NCBI.

bioinformatics↗

Integrated population clustering and genomic epidemiology with PopPIPE

Genetic distances between bacterial DNA sequences can be used to cluster populations into closely related subpopulations, and as an additional source of information when detecting possible transmission events. Due to their variable gene content and order, reference-free methods offer more sensitive detection of genetic differences, especially among closely related samples found in outbreaks. However, across longer genetic distances, frequent recombination can make calculation and interpretation of these differences more challenging, requiring significant bioinformatic expertise and manual intervention during the analysis process. Here we present a Population analysis PIPEline (PopPIPE) which combines rapid reference-free genome analysis methods to analyse bacterial genomes across these two scales, splitting whole populations into subclusters and detecting plausible transmission events within closely related clusters. We use k-mer sketching to split populations into strains, followed by split k-mer analysis and recombination removal to create alignments and subclusters within these strains. We first show that this approach creates high quality subclusters on a population-wide dataset of Streptococcus pneumoniae. When applied to nosocomial vancomycin resistant Enterococcus faecium samples, PopPIPE finds transmission clusters which are more epidemiologically plausible than core genome or MLST-based approaches. Our pipeline is rapid and reproducible, creates interactive visualisations, and can easily be reconfigured and re-run on new datasets. Therefore PopPIPE provides a user-friendly pipeline for analyses spanning species-wide clustering to outbreak investigations. Impact statementAs time passes, bacterial genomes accumulate small changes in their sequence due to mutations, or larger changes in their content due to horizontal gene transfer. Using their genome sequences, it is possible to use phylogenetics to work out the most likely order in which these changes happened, and how long they took to happen. Then, one can estimate the time that separates any two bacterial samples - if it is short then they may have been directly transmitted or acquired from the same source; but if it is long they must have been acquired separately. This information can be used to determine transmission chains, in conjunction with dates and locations of infections. Understanding transmission chains enables targeted infection control measures. However, correctly calculating the genetic evidence for transmission is made difficult by correctly distinguishing different types of sequence changes, dealing with large amounts of genome data, and the need to use multiple complex bioinformatic tools. We addressed this gap by creating a computational workflow, PopPIPE, which automates the process of detecting possible transmissions using genome sequences. PopPIPE applies state-of-the-art tools and is fast and easy to run - making this technology will be available to a wider audience of researchers. Data summaryThe code for this pipeline is available at https://github.com/bacpop/PopPIPE and as a docker image https://hub.docker.com/r/poppunk/poppipe. Raw sequencing reads for Enterococcus faecium isolates have been deposited at the NCBI under BioProject accession number PRJNA997588.

bioinformatics↗