bioRxiv Science⌕ Search

Biology subjects

Horsfield, S. T.

Publications and source records attributed to Horsfield, S. T..

9 recordsLinked to original sources

Species-specific transformer models of bacterial gene order and content for genomic surveillance tasks

Transformer models enable functionally meaningful representation of complex biological data, such as nucleotide or protein sequences. Existing foundation transformer models are trained on large multi-domain corpuses of unlabelled DNA or protein data, showing unmatched task generalisation. However, these foundation models are often outperformed on domain-specific tasks by models trained on taxonomically-constrained data, such as prokaryote gene annotation. By extension, species-specific transformer models hold promise for targeted analyses, given sufficient training data are available. Epidemiological analysis of bacterial pathogens exemplifies the use case of species-specific transformers, due to the wealth of genome data available, coupled with pathogen-specific analyses carried out during routine and outbreak surveillance. Here, we trained a transformer model, PanBART, on the gene content and gene order of two important and biologically distinct bacterial pathogens, Escherichia coli and Streptococcus pneumoniae, benchmarking against state-of-the-art non-transformer approaches for genomic epidemiology. We show PanBART learns representations of population structure in an unsupervised manner, and can be used to accurately assign genomes to biologically-meaningful sequence clusters. PanBART is also able to identify emergent lineages, differentiating them from pre-existing lineages, and can accurately predict genomes likely to uptake genes involved in antibiotic resistance before a transfer event has occurred. Finally, PanBART can be used to conduct co-selection analysis to identify pairs of genes likely to be evolving together. Our work demonstrates that species-specific transformer models can be employed in many critical public health scenarios. We lay the groundwork for wider application of such models in epidemiological analysis, and provide scenarios where such models excel.

bioinformatics↗

Rapid gene exchange explains differences in bacterial pangenome structure

The size and diversity of bacterial gene repertoires, known as pangenomes, vary widely across species. The evolutionary forces driving the maintenance of pangenomes is an open topic of debate, with contradictory theories suggesting that pangenomes exist as a result of neutral evolution, with all genes gained and lost at random, or that all genes provide a fitness benefit to the host and are maintained by positive selection. Modelling of pangenome dynamics has provided insight into how gene exchange explains observed gene frequency distributions, and stands as the only means of jointly inferring contributions of individual gene selection effects and mobility on the maintenance of pangenomes. However, previous modelling studies have not included both gene-level selection and mobility, and do not consider broadly sampled genome datasets for many species. To differentiate neutral and selective forces maintaining pangenomes, we developed a mechanistic model of gene-level evolution, Pansim, and a scalable model fitting framework, PopPUNK-mod. Together, these tools leverage rapid genome distance calculation to fit models of pangenome dynamics to datasets containing hundreds of thousands of genomes. We used this framework to compare the pangenome dynamics of over 400 different bacterial species, using over 600,000 genomes. We find that diversity in pangenome characteristics between species is driven predominantly by variation in the number of rapidly exchanged genes, while the rate of exchange of remaining genes is conserved. We find that bacterial phylogeny, rather than ecology, correlates with pangenome dynamics. We express that pan-species gene-level analyses are now needed to understand selection across accessory genes. Our work highlights the importance of gene exchange rate differences in governing differences in pangenome characteristics between species.

bioinformatics↗

A reusable model of pangenome selection informs optimal surveillance strategies over vaccine introductions

BackgroundThe human pathogen Streptococcus pneumoniae is a major cause of disease, including pneumonia and meningitis. The introduction of Pneumococcal Conjugate Vaccines (PCVs) initially reduced the burden of disease through a reduction of colonisation by vaccine-targeted serotypes. However, since PCVs only target a proportion of pneumococcal serotypes, they shift intraspecific competition, eventually allowing non-targeted types to replace vaccine types. Understanding the host and pathogen factors causing replacement is important for future vaccine development. Mechanistic understanding of vaccine replacement dynamics is crucial for forecasting and optimisation of genomic surveillance strategies to evaluate realised vaccine effectiveness. MethodsWe developed a mathematical model of the genomic and demographic factors which explain vaccine replacement, used this model to replicate serotype-frequency changes, and investigated cost-effective genomic surveillance strategies. We extended a forward-time model based on the Wright-Fisher model, developing a user-friendly model framework that describes the post-vaccine dynamics of S. pneumoniae populations. Our model describes vaccine replacement as a function of vaccine impact, immigration of new strains, and negative frequency-dependent selection (NFDS) on the accessory genome content. ResultsWe used our model to study vaccine replacement in newly sequenced genomic surveillance data from Nepal, and existing data from the US, and the UK, with distinct surveillance strategies. We showed that the model with NFDS better replicates replacement dynamics than a null model without NFDS, and that NFDS likely only acts on part of the S. pneumoniae accessory genome. We found consistent estimates for vaccination effectiveness across the different study locations and country-specific genes under NFDS, highlighting the importance of conducting genomic surveillance in each country of interest. By simulating data from the model, we showed that an optimal surveillance strategy prioritises per-sampling sample size over sampling frequency for small sampling budgets. ConclusionsOur model can be used to predict vaccine replacement dynamics after PCV introduction, and can be easily reapplied to analyse new data from vaccine introductions or new regions. Our model is available in the R package STUBENTIGER (Studying Balancing Evolution (NFDS) To Investigate Genome Replacement) on GitHub https://github.com/bacpop/Stubentiger.

genomics↗

Selection-free CRISPR-Cas9 editing protocol for distant Dictyostelid species

Dictyostelids are a species-rich clade of cellular slime molds that are widely found in soils and have been studied for over a century. Most research focusses on Dictyostelium discoideum, which - due to its ease of culturing and genetic tractability - has been adopted as a model species in the fields of developmental biology, cell biology and microbiology. Over decades, genome editing methods in D. discoideum have steadily improved but remain relatively time-consuming and limited in scope, effective in a few species only. Here, we introduce a CRISPR-Cas9 editing protocol that is cloning-free, selection-free, highly-efficient, and effective across Dictyostelid species. After optimizing our protocol in D. discoideum, we obtained knock-out efficiencies of [~]80% and knock-in efficiencies of [~]30% without antibiotic selection. Efficiencies depend on template concentrations, insertion sizes, homology arms and target sites. Since our protocol is selection-free, we can isolate mutants as soon as one day post-transfection, vastly expediting the generation of knock-outs, fusion proteins and expression reporters. Our protocol also makes it possible to generate several knock-in mutations simultaneously in the same cells. Boosted by cell-sorting and fluorescent microscopy, we could readily apply our CRISPR-Cas9 editing protocol to phylogenetically distant Dictyostelid species, which diverged hundreds of millions of years ago and have never been genome edited before. Our protocol therefore opens the door to performing broad-scale genetic interrogations across Dicyostelids.

genetics↗

Pneumococcal chromosome conformation variation between epigenetic variants is driven by episomal mobile genetic elements

Streptococcus pneumoniae (pneumococcus) is a genetically diverse opportunistic bacterial pathogen that expresses two phase-variable loci encoding restriction-modification systems. Comparisons of two genetically-distinct pairs of epigenetically-distinct variants, each distinguished by a stabilised arrangement of one of these phase-variable loci, found the consequent changes in genome-wide DNA methylation patterns were associated with differential expression of mobile genetic elements (MGEs). This relationship was hypothesised to be mediated through changes in xenogenic silencing (XS) or nucleoid organisation. Therefore the chromosomal conformation of the both variants of each isolate were characterised using Illumina Hi-C, and Nanopore Pore-C, sequencing. Both methods concurred that the organisation of the pneumococcal chromosome was dominated by small-scale structures, with most pairwise interactions between loci <25 kb apart. Neither found substantial evidence for higher-order structure or XS in the pneumococcal genome, with more complex contact patterns only evident around the replication origin. Comparisons between the variants identified phage-related chromosomal islands (PRCIs) as the foci of differential contact densities between the variants. This was driven by copy number variation, resulting from variable excision and replication of the episomal PRCIs. However, the methods were discordant in their identification of the variant in which the PRCI was more actively replicating in both pairs. Validatory experiments demonstrated that the prevalence of circular PRCIs was not determined by DNA modification, but instead varied stochastically between colonies in both backgrounds, and was metastable during vegetative growth. PRCI excision was inducible by mitomycin C, but independent of the presence of a phage. Yet transcriptional activation of these elements was affected by both signals, indicating transcription and replication are separately regulated. Therefore pneumococcal MGEs do not appear to be subject to XS, resulting in heterogeneity being generated within these bacterial populations through the frequent local disruption of chromosome conformation resulting from the stochastic excision and reintegration of episomal elements. Author summaryThe pneumococcus is a bacterium with a circular chromosome that is organised by DNA-binding proteins and often contains mobile genetic elements (MGEs), genes able to transmit between bacteria. All pneumococci encode defences against MGEs, some of which create epigenetic modifications (typically methylation) genome-wide at particular sequence motifs. Changes in these epigenetic patterns are associated with altered MGE gene expression and replication in otherwise genetically-identical bacteria. To test whether this was the result of methylation remodelling the organisation of the chromosome, we compared the contact patterns across the chromosome using two sequencing technologies. Both methods concurred that the pneumococcal genome is generally folded into small structures, with the biggest differences between the variants caused by the replication of MGEs. However, the methods disagreed on the variant in which the MGEs replicated fastest. Further experiments showed that MGE replication was stable over the course of culturing over hours, but would randomly change level between days, explaining the inconsistent observations. MGE replication was found to rise in response to DNA damage, whereas gene expression also depended on the presence of other signals, explaining the discrepancies in these activities between variants. Hence MGEs significantly contribute to the heterogeneity that rapidly accumulates within pneumococcal populations.

genomics↗

Integrated population clustering and genomic epidemiology with PopPIPE

Genetic distances between bacterial DNA sequences can be used to cluster populations into closely related subpopulations, and as an additional source of information when detecting possible transmission events. Due to their variable gene content and order, reference-free methods offer more sensitive detection of genetic differences, especially among closely related samples found in outbreaks. However, across longer genetic distances, frequent recombination can make calculation and interpretation of these differences more challenging, requiring significant bioinformatic expertise and manual intervention during the analysis process. Here we present a Population analysis PIPEline (PopPIPE) which combines rapid reference-free genome analysis methods to analyse bacterial genomes across these two scales, splitting whole populations into subclusters and detecting plausible transmission events within closely related clusters. We use k-mer sketching to split populations into strains, followed by split k-mer analysis and recombination removal to create alignments and subclusters within these strains. We first show that this approach creates high quality subclusters on a population-wide dataset of Streptococcus pneumoniae. When applied to nosocomial vancomycin resistant Enterococcus faecium samples, PopPIPE finds transmission clusters which are more epidemiologically plausible than core genome or MLST-based approaches. Our pipeline is rapid and reproducible, creates interactive visualisations, and can easily be reconfigured and re-run on new datasets. Therefore PopPIPE provides a user-friendly pipeline for analyses spanning species-wide clustering to outbreak investigations. Impact statementAs time passes, bacterial genomes accumulate small changes in their sequence due to mutations, or larger changes in their content due to horizontal gene transfer. Using their genome sequences, it is possible to use phylogenetics to work out the most likely order in which these changes happened, and how long they took to happen. Then, one can estimate the time that separates any two bacterial samples - if it is short then they may have been directly transmitted or acquired from the same source; but if it is long they must have been acquired separately. This information can be used to determine transmission chains, in conjunction with dates and locations of infections. Understanding transmission chains enables targeted infection control measures. However, correctly calculating the genetic evidence for transmission is made difficult by correctly distinguishing different types of sequence changes, dealing with large amounts of genome data, and the need to use multiple complex bioinformatic tools. We addressed this gap by creating a computational workflow, PopPIPE, which automates the process of detecting possible transmissions using genome sequences. PopPIPE applies state-of-the-art tools and is fast and easy to run - making this technology will be available to a wider audience of researchers. Data summaryThe code for this pipeline is available at https://github.com/bacpop/PopPIPE and as a docker image https://hub.docker.com/r/poppunk/poppipe. Raw sequencing reads for Enterococcus faecium isolates have been deposited at the NCBI under BioProject accession number PRJNA997588.

bioinformatics↗

CELEBRIMBOR: Pangenomes from metagenomes

SummaryMetagenome Assembled Genomes (MAGs) are often incomplete, with sequences missing due to errors in assembly or low coverage. Incomplete MAGs present a particular challenge for identification of shared genes within a microbial population, known as core genes, as a core gene missing in only a few assemblies will result in it being mischaracterized at a lower frequency. Here, we present CELEBRIMBOR, a snakemake pangenome analysis pipeline which uses a measure of genome completeness to automatically adjust the frequency threshold at which core genes are identified, enabling accurate core gene identification in MAGs. Availability and implementationCELEBRIMBOR is published under open source Apache 2.0 licence at https://github.com/bacpop/CELEBRIMBOR and is available as a Docker container. Supplementary material is available in the online version of the article.

bioinformatics↗

Graph-based Nanopore Adaptive Sampling with GNASTy enables sensitive pneumococcal serotyping in complex samples

Serotype surveillance of Streptococcus pneumoniae (the pneumococcus) is critical for understanding the effectiveness of current vaccination strategies. However, existing methods for serotyping are limited in their ability to identify co-carriage of multiple pneumococci and detect novel serotypes. To develop a scalable and portable serotyping method that overcomes these challenges, we employed Nanopore Adaptive Sampling (NAS), an on-sequencer enrichment method which selects for target DNA in real-time, for direct detection of S. pneumoniae in complex samples. Whereas NAS targeting the whole S. pneumoniae genome was ineffective in the presence of non-pathogenic streptococci, the method was both specific and sensitive when targeting the capsular biosynthetic locus (CBL), the operon that determines S. pneumoniae serotype. NAS significantly improved coverage and yield of the CBL relative to sequencing without NAS, and accurately quantified the relative prevalence of serotypes in samples representing co-carriage. To maximise the sensitivity of NAS to detect novel serotypes, we developed and benchmarked a new pangenome-graph algorithm, named GNASTy. We show that GNASTy outperforms the current NAS implementation, which is based on linear genome alignment, when a sample contains a serotype absent from the database of targeted sequences. The methods developed in this work provide an improved approach for novel serotype discovery and routine S. pneumoniae surveillance that is fast, accurate and feasible in low resource settings. GNASTy therefore has the potential to increase the density and coverage of global pneumococcal surveillance. One sentence summaryPangenome graph-based Nanopore Adaptive Sampling, presented in our tool GNASTy, is a sensitive, portable and cost-effective method for Streptococcus pneumoniae surveillance.

genomics↗

Accurate and fast graph-based pangenome annotation and clustering with ggCaller

Bacterial genomes differ in both gene content and sequence mutations, which can cause important clinical phenotypic differences such as vaccine escape or antimicrobial resistance. To identify and quantify important variants, all genes within a population must be predicted, functionally annotated and clustered, representing the pangenome. Despite the volume of genome data available, gene prediction and annotation are currently conducted in isolation on individual genomes, which is computationally inefficient and frequently inconsistent across genomes. Here, we introduce the open-source software graph-gene-caller (ggCaller; https://github.com/samhorsfield96/ggCaller). ggCaller combines gene prediction, functional annotation and clustering into a single step using population-wide de Bruijn Graphs, removing redundancy in gene annotation, and resulting in more accurate gene predictions and orthologue clustering. We applied ggCaller to simulated and real-world bacterial genome datasets, comparing it to current state-of-the-art tools. ggCaller is ~50x faster with equivalent or greater accuracy, particularly in datasets with complex sources of error, such as assembly contamination or fragmentation. ggCaller is also an important extension to bacterial genome-wide association studies, enabling querying of annotated graphs for functional analyses. We highlight this application by functionally annotating DNA sequences with significant associations to tetracycline and macrolide resistance in Streptococcus pneumoniae, identifying key resistance determinants that were missed when using only a single reference genome. ggCaller is a novel bacterial genome analysis tool with applications in bacterial epidemiology and evolutionary study.

bioinformatics↗