bioRxiv Science⌕ Search

Biology subjects

Parfitt, K. M.

Publications and source records attributed to Parfitt, K. M..

5 recordsLinked to original sources

High resolution Streptococcus pyogenes core genome MLST and LIN coding scheme for outbreak detection

Streptococcus pyogenes is a globally important pathogen responsible for at least 500,000 deaths a year, causing significant burden on healthcare systems. It is the causative agent for ailments such as impetigo and strep throat to septicaemia and necrotizing fasciitis. Assessment of genetic relatedness for the detection of outbreaks within communities or healthcare facilities is vital in decreasing the propagation of S. pyogenes within these settings, alongside epidemiological data. As the volume of isolates being sequenced increases year on year, more scalable and sharable methodologies of assessing genetic relatedness are required by reference laboratories and for international collaboration. LIN codes, applied to core genome MLST (cgMLST) represent a method which is extensible to large scale whole genome sequencing (WGS) while still being sufficiently sensitive to detect outbreak clusters. Here we present a novel cgMLST and LIN code scheme, hosted by PubMLST, enabling international collaboration and global tracking of variants, that is highly scalable and usable for all. The schemes are available at https://pubmlst.org/organisms/streptococcus-pyogenes. Data SummaryGenome sequences and metadata are available at https://pubmlst.org/organisms/streptococcus-pyogenes. PubMLST and ENA accessions and metadata can additionally be found in the supplementary data. Raw reads for UKHSA sequences are available in ENA study PRJEB115996. Impact StatementStreptococcus pyogenes is a globally relevant pathogen capable of causing invasive and non-invasive disease across a multitude of settings. Assessment of genetic relatedness is an increasingly important aspect of managing outbreaks, requiring solutions that are scalable, high resolution and comparable across laboratories. Here we present a high resolution core genome multi locus sequence typing (MLST) and associated life identification number (LIN) code scheme, The schemes were developed using a combination of 4,916 UKHSA and 2,391 publicly available S. pyogenes isolates in order to cover a wide range of EMM types both within the UK and globally. These new schemes enable high resolution typing of S. pyogenes isolates, suitable for analysis of lineages to genomic epidemiology in outbreak detection and management. Both cgMLST and LIN code schemes are available on PubMLST as an open access resource for the public health and academic communities and can enable both intra laboratory and global coordination.

bioinformatics↗

A Life Identification Number Barcoding (LIN Code) System for Neisseria meningitidis: high resolution multi-level typing of meningococci.

Neisseria meningitidis is a commensal member of the human oropharyngeal microbiota that can cause devastating invasive meningococcal disease. This genetically and antigenically diverse accidental pathogen has been a paradigm for the study of bacterial population biology. The meningococcus was the first organism for which a seven-locus multi-locus sequence typing (MLST) scheme was developed. With the addition of sequence-based characterisation of antigen genes, molecular typing has been widely employed to inform surveillance and public health interventions. Following the advent and widespread adoption of whole genome sequencing (WGS), precise delineation of variants is possible. Here, a WGS-based Life Identification Number (LIN) code typing scheme is described, providing a multi-resolution nomenclature for understanding meningococcal population diversity and molecular epidemiology. The LIN codes were developed using a set of 6,131 N. meningitidis genomes, comprising up to 200 isolates from each clonal complex (cc) previously described using MLST. Based on cluster-stability analysis and concordance with ccs, thirteen LIN thresholds described the meningococcal population at different levels of resolution. LIN codes and human-readable nicknames consistent with existent nomenclatures were assigned to the N. meningitidis genomes hosted in the PubMLST. Published outbreaks validated the LIN thresholds, illustrating the potential of this genomic tool in public health management.

microbiology↗

A stable, hierarchical LIN code system for Campylobacter jejuni and Campylobacter coli: A unified genomic nomenclature for lineage-level typing and global surveillance.

Campylobacter remains the leading cause of bacterial gastroenteritis worldwide, with C. jejuni accounting for around 90% of infection and C. coli accounting for most of the rest. Seven-locus multilocus sequence typing (MLST) has improved our understanding of host association and population structure, whilst core genome MLST (cgMLST), enables investigation of transmission events at high-resolution. However, the lack of a stable and standardised nomenclature for clustering of cgMLST data has limited reproducibility and long-term comparability between studies. Here we introduce a joint, hierarchical Life Identification Number (LIN) code system that provides reproducible, multi-level genomic identifiers for C. jejuni and C. coli lineages. Using an updated cgMLST v2 scheme (1,142 loci) and globally representative datasets of high-quality genomes selected from over 53,000 assemblies in the Campylobacter PubMLST database (https://pubmlst.org/organisms/campylobacter-jejunicoli), we firstly defined LIN codes on a dataset of 5,664 genomes. Pairwise allelic distances were computed using MSTclust, and 18 nested thresholds were defined through silhouette, adjusted Wallace and adjusted Rand Index (ARI) statistics to capture the population structure from species to outbreak level resolution. The LIN thresholds were then validated using a second dataset of 1,781 genomes from PubMLST and applied to a large water-associated outbreak dataset from New Zealand in 2016, containing clinical and ecological genomes. Further application of LIN codes was demonstrated by analyses of the C. jejuni ST-21 clonal complex and ST-6175 isolates, as well as the broader population structure of C. coli, using data from PubMLST. Across all datasets, LIN clusters were stable, largely monophyletic, and back-compatible with existing nomenclature, accurately distinguishing host-adapted and outbreak-associated lineages. By embedding cgMLST data within a stable and scalable nomenclature, the Campylobacter LIN system delivers consistent, automated genome-to-lineage assignment. This unified framework bridges population genetics and applied surveillance, enabling robust, real-time comparison of Campylobacter isolates across sources, studies, and time. Impact statementHuman cases of Campylobacter worldwide continue unabated. Tracing the source of Campylobacter infection is particularly challenging given the sporadic or multi-source nature of outbreaks, with potential transmission from foodborne, animal or environmental sources. Seven-locus MLST has greatly improved our broad understanding of Campylobacter population structure. However, whilst high-resolution cgMLST alleles and STs themselves do not change, longitudinal cluster analyses of cgMLST data have lacked a stable nomenclature, rendering them unsuitable for robust and comparable surveillance over time. Life Identification Number (LIN) codes provide a solution to this problem, establishing an automated and scalable nomenclature derived directly from cgMLST profiles, that is stable over time. We have implemented a joint C. jejuni and C. coli LIN code scheme in PubMLST, with scripts for real-time lineage assignment. LIN codes are back-compatible with existing MLST nomenclature, and we demonstrate their added practical value for exploring population structure and high-resolution outbreak investigation. LIN codes support surveillance of Campylobacter in a One Health context, by enabling consistent typing at multiple levels across different sources, laboratories and time. Data summary1. The isolate collections used to develop the LIN codes are publicly available and searchable as individual projects on the PubMLST database (https://pubmlst.org). O_LILIN code development (Dataset 1) (n=5,664 isolates, up to 200 isolates per clonal complex) C_LIO_LILIN code validation (Dataset 2) (n=1,781 isolates), up to 50 isolates per clonal complex C_LIO_LIOutbreak investigation (Dataset 3): New Zealand 2016 Havelock North waterborne outbreak, Gilpin et al (n=161 isolates) [1] C_LIO_LIPopulation structure exploration (clonal complex) (Dataset 4): ST-21 complex (n=1800 isolates, up to 100 isolates randomly selected from each country) C_LIO_LIPopulation structure exploration (sequence type) (Dataset 5): ST-6175 (n=321 isolates, genomes with good cgMLST v2 annotation) C_LI 2. The software for LIN code development is publicly available as follows: O_LIMSTclust for pairwise distance matrices https://gitlab.pasteur.fr/GIPhy/MSTclust [2] C_LIO_LIPython script to define LIN codes in a local dataset; (https://gitlab.pasteur.fr/BEBP/LINcoding) C_LIO_LIBIGSdb Perl script to define LIN codes from cgMLST profiles on the PubMLST database; (https://github.com/kjolley/BIGSdb/blob/develop/scripts/maintenance/lincodes.pl) C_LI

microbiology↗

Neisseria gonorrhoeae LIN codes: a Robust, Multi-Resolution Lineage Nomenclature

Investigation of the bacterial pathogen Neisseria gonorrhoeae is complicated by extensive horizontal gene transfer: a process which disrupts phylogenetic signals and impedes our understanding of population structure. The ability to consistently identify N. gonorrhoeae lineages is important for surveillance of this increasingly antimicrobial resistant organism, facilitating efficient communication regarding its epidemiology; however, conventional typing systems fail to reflect N. gonorrhoeae strain taxonomy in a reliable and stable manner. Here, a N. gonorrhoeae genomic lineage nomenclature, based on the barcoding system of Life Identification Number (LIN) codes, was developed using a refined 1430 core gene MLST (cgMLST). This hierarchical LIN code nomenclature conveys lineage information at multiple levels of resolution within one code, enabling it to provide immediate context to an isolates ancestry, and to relate to familiar, previously used typing schemes such as Ng cgMLST v1, 7-locus MLST, or NG-STAR clonal complex (CC). Clustering with LIN codes accurately reflects gonococcal diversity and population structure, providing insight into associations between genotype and phenotype for traits such as antibiotic resistance. These codes are automatically assigned and publicly accessible via the pubmlst.org/organisms/neisseria-spp database.

genomics↗

Identification of two distinct phylogenomic lineages and model strains for the understudied cystic fibrosis lung pathogen Burkholderia multivorans

Burkholderia multivorans is the dominant Burkholderia pathogen recovered from lung infection in people with cystic fibrosis. However, as an understudied pathogen there are knowledge gaps in relation to its population biology, phenotypic traits and useful model strains. A phylogenomic study of B. multivorans was undertaken using a total of 283 genomes, of which 73 were sequenced and 49 phenotypically characterized as part of this study. Average nucleotide identity analysis (ANI) and phylogenetic alignment of core genes demonstrated that the B. multivorans population separated into two distinct evolutionary clades, defined as lineage 1 (n = 58 genomes) and lineage 2 (n = 221 genomes). To examine the population biology of B. multivorans, a representative subgroup of 77 B. multivorans genomes (28 from the reference databases and the 49-novel short-read genome sequences) were selected based on multilocus sequence typing (MLST), isolation source and phylogenetic placement criteria. Comparative genomics was used to identify B. multivorans lineage-specific genes: ghrB_1 in lineage 1, and glnM_2 in lineage 2, and diagnostic PCRs targeting them successfully developed. Phenotypic analysis of 49 representative B. multivorans strains showed considerable variance with the majority of isolates tested being motile and capable of biofilm formation. A striking absence of B. multivorans protease activity in vitro was observed, but no lineage-specific phenotypic differences demonstrated. Using phylogenomic and phenotypic criteria, three model B. multivorans CF strains were identified, BCC0084 (lineage 1), BCC1272 (lineage 2a) and BCC0033 lineage 2b, and their complete genome sequences determined. B. multivorans CF strains BCC0033 and BCC0084, and the environmental reference strain, ATCC 17616, were all capable of short-term survival within a murine lung infection model. By mapping the population biology, identifying lineage-specific PCRs and model strains, we provide much needed baseline resources for future studies of B. multivorans.

microbiology↗