bioRxiv Science⌕ Search

Biology subjects

Rethoret-Pasty, M.

Publications and source records attributed to Rethoret-Pasty, M..

7 recordsLinked to original sources

Cumulative cgMLST provides increased discrimination of nested phylogenetic groups

BackgroundCore genome multilocus sequence typing (cgMLST) is a powerful method for bacterial strain genotyping. However, the size of the core genome decreases as the phylogenetic breadth of the target group increases, reducing discriminatory power. To overcome this discrimination/applicability tradeoff, here we developed a cumulative cgMLST approach, where sets of core loci conserved within nested phylogenetic entities are added. We illustrate this approach using the Klebsiella pneumoniae species complex (KpSC), for which a widely used cgMLST scheme (KpSC-cgMLST) comprises only 629 genes. MethodsWe created non-redundant cgMLST schemes for the individual species K. pneumoniae sensu stricto (Kpn-cgMLST scheme), and its multidrug resistant sublineages (SLs) SL147 and SL307. To extract core genes, we used 37,874 genome assemblies originating from over 80 countries worldwide. A methodology was set to filter redundant loci before importing them into the genotyping tool BIGSdb, where they were combined into schemes together with preexisting loci conserved at higher phylogenetic levels. The performance of the cumulative cgMLST schemes was evaluated on previously published datasets and on novel data from an inter-hospital outbreak of SL307. ResultsThe Kpn-cgMLST, SL147 and SL307 schemes comprise 2752, 852, and 947 additional loci, respectively. The mean allele call rate of the novel loci was >99% in the validation datasets. Compared to the KpSC scheme used alone, pairwise allelic distances among isolates increased on average 5.6-fold using the Kpn scheme, and further by 20% and 30% using the SL147 and SL307 schemes, respectively. We demonstrate the added value of this increased discriminatory power for epidemiological analyses and show nearly equal discrimination when compared to whole-genome single nucleotide polymorphisms analysis. ConclusionsThe cumulative cgMLST strategy combines broad phylogenetic applicability and nearly complete genotyping resolution, expanding the utility of this harmonized approach for genomic epidemiology.

microbiology↗

Life identification number (LIN) codes for the genomic taxonomy of Corynebacterium diphtheriae strains

BackgroundCorynebacterium diphtheriae, which causes diphtheria, remains a public health concern especially in regions with low vaccination coverage. While advances in genomic typing, such as core-genome Multi-Locus Sequence Typing (cgMLST, based on 1305 genes), have improved our ability for strain identification, a standardized and stable genomic taxonomy is still lacking. This study aimed to establish a consistent classification and nomenclature for C. diphtheriae strains using cgMLST-based Life Identification Number (LIN) codes. MethodsComparing 1,665 genomes from C. diphtheriae and its closely related species C. belfantii and C. rouxii, we observed population-level genetic discontinuities in cgMLST profiles dissimilarities, and established hierarchical taxonomic levels based on optimal allelic difference thresholds. Ten-level LIN codes were defined, encompassing broad population structure subdivisions and fine-scale epidemiological levels. The LIN code system was implemented into the BIGSdb-Pasteur platform, and nicknames derived from the 7-loci MLST sequence types were given to sublineages and clonal groups. ResultscgMLST genetic thresholds were first defined at species (minimum of 1,240 allelic differences) and lineage levels (1,035 differences). Sublineages (SL), clonal groups (ClG), and genetic clusters (GC) were next defined with progressively finer allelic mismatch thresholds (500, 55, and 25 differences, respectively). A broad population diversity of C. diphtheriae was uncovered, with the distinction of >400 SLs and >1,000 GCs. For epidemiological purposes, five shallow-level thresholds (8, 4, 2 ,1, and 0 allelic mismatches were defined, completing the 10-level LIN code taxonomy. We illustrate LIN codes applicability to investigate the genetic diversity and transmission chains of relevant clusters, such as SL8 (the 1990s ex-USSR outbreak) or SL384 (involved in outbreaks in Yemen and Europe). ConclusionsThe cgMLST-based LIN code system provides a stable genomic taxonomy for strains of C. diphtheriae, C. rouxii and C. belfantii. By defining ten hierarchical levels of resolution, this system effectively captures its phylogenetic diversity, facilitating population biology research and epidemiological surveillance. The public availability of this system from the BIGSdb-Pasteur platform provides a standardized framework for diphtheria genomic epidemiology with potential to harmonize global surveillance of the resurgence of diphtheria.

microbiology↗

Accurate genotyping of three major respiratory bacterial pathogens with ONT R10.4.1 long-read sequencing

High-throughput massive parallel sequencing has significantly improved bacterial pathogen genomics, diagnostics, and epidemiology. Despite its high accuracy, short-read sequencing struggles with complete genome reconstruction and assembly of extrachromosomal elements such as plasmids. Long-read sequencing with Oxford Nanopore Technologies (ONT) presents an alternative that offers benefits like real-time sequencing and cost-efficiency, particularly useful in resource-limited settings. However, the higher error rates of ONT have so far limited its application in high-precision genomic typing. The recent release of ONTs R10.4.1 chemistry, with significantly improved raw read accuracy (Q20+), offers a potential solution to this problem. The aim of this study was to evaluate the performance of ONTs latest chemistry for bacterial genomic typing against the gold standard Illumina technology, focusing on three respiratory pathogens of public health importance, Klebsiella pneumoniae, Bordetella pertussis, and Corynebacterium diphtheriae, and their related species. Using the Rapid Barcoding Kit V14, we generated and analyzed genome assemblies with different basecalling tools and models, at different simulated depths of coverage. ONT assemblies were compared to the Illumina reference for completeness and core genome multilocus sequence typing (cgMLST) accuracy (number of allelic mismatches). Our results show that genomes obtained from raw data basecalled with Dorado (with both simplex and duplex reads) SUP v0.7.1, assembled with Flye, and with a minimum coverage depth of 30x, optimized the accuracy for all bacterial species tested. The error rates were consistently below 1% of each cgMLST scheme, indicating that ONT R10.4.1 data is suitable for high-resolution genomic typing applied to outbreak investigations and public health surveillance.

genomics↗

Genomic Epidemiology and Microevolution of the Zoonotic Pathogen Corynebacterium ulcerans

Corynebacterium ulcerans is an emerging zoonotic pathogen that belongs to the Corynebacterium diphtheriae (Cd) species complex (CdSC), and that causes diphtheria-like infections in humans. Our understanding of the transmission, phylogeography and evolution of C. ulcerans remains limited, in part due to the lack of a standardized genomic epidemiology toolkit. The aim of this work was to develop a core genome multi-locus sequence typing (cgMLST) scheme for high-resolution genotyping and classification of C. ulcerans strains, and to explore transmission, spatial spread and genomic evolution among 582 C. ulcerans isolates from sporadic clinical cases and reported case clusters. The cgMLST scheme combines 1,628 loci with highly reproducible allele calls and shows high strain subtyping resolution. We demonstrate its utility for capturing population structure by defining sublineages (SL, maximum 940 allele differences) and clonal groups (CG, 194 allele differences, AD) and for epidemiological surveillance by defining genetic clusters, i.e., previously undetected chains of transmission (25 AD). Genetic clusters correspond to cryptic and case clusters that were associated with specific geographical regions within France. Major C. ulcerans sublineages (SL325, SL331, SL339) and clonal groups (CG325, CG331, CG583) showed strong associations with diphtheria toxin variants and tox-carrying prophages or other genetic elements. The evolutionary dynamics of tox gene presence or absence varied sharply among clonal groups. The cgMLST scheme is publicly available (https://bigsdb.pasteur.fr/diphtheria) and provides a common framework for investigating the ecology, evolution and variations in virulence among C. ulcerans strains. The implementation of a standardized high-resolution genotyping method will also facilitate the tracing of C. ulcerans transmission and spread across hosts and from local to global spatial scales.

microbiology↗

A metabolic atlas of the Klebsiella pneumoniae species complex reveals lineage-specific metabolism that supports persistent co-existence of diverse lineages

The Klebsiella pneumoniae species complex inhabits a wide variety of hosts and environments, and is a major cause of antimicrobial resistant infections. Genomics has revealed the population comprises multiple species/subspecies and hundreds of distinct co-circulating sub-lineages that are associated with distinct gene complements. A substantial fraction of the pan-genome is predicted to be involved in metabolic functions and hence these data are consistent with metabolic differentiation as a driver of population structure. However, this has so far remained unsubstantiated because in the past it was not possible to explore metabolic variation at scale. Here we used a combination of comparative genomics and high-throughput genome-scale metabolic modelling to systematically explore metabolic diversity across the K. pneumoniae species complex (n=7,835 genomes). We simulated growth outcomes for each isolate using carbon, nitrogen, phosphorus and sulfur sources under aerobic and anaerobic conditions (n=1,278 conditions per isolate). We showed that the distributions of metabolic genes and growth capabilities are structured in the population, and confirmed that sub-lineages exhibit unique metabolic profiles. In vitro co-culture experiments demonstrated reciprocal commensalistic cross-feeding between sub-lineages, effectively extending the range of conditions supporting individual growth. We propose that these substrate specialisations promote the existence and persistence of co-circulating sub-lineages by reducing nutrient competition and facilitating commensal interactions via negative frequency-dependent selection. Our findings have implications for understanding the eco-evolutionary dynamics of K. pneumoniae and for the design of novel strategies to prevent opportunistic infections caused by this World Health Organization priority antimicrobial resistant pathogen.

microbiology↗

Bacterial strain nomenclature in the genomic era: Life Identification Numbers using a gene-by-gene approach

Unified strain taxonomies are needed for the epidemiological surveillance of bacterial pathogens and international communication in microbiological research. Core genome multilocus sequence typing (cgMLST) holds great promise for standardized high-resolution strain genotyping. However, this approach faces challenges including classification instability and disconnection of new nomenclature from widely adopted classical MLST identifiers. This essay discusses the cgMLST-based Life Identification Number (LIN) method, recently proposed as a stable multilevel strain taxonomy system applicable to most bacterial pathogens. We describe how LIN codes are implemented and used in practice for precise strain definitions and epidemiological tracking. Glossary Multilocus sequence typing (MLST)A genotyping method applied mostly to microbial strains to study population structure and epidemiology, based on comparing the nucleotide sequences of a small number (typically seven) of housekeeping protein-coding genes. In MLST, allele numbers are assigned to each sequence variant (allele) of a given gene. The MLST genotype of a bacterial strain is defined by the combination of the allele numbers observed at the genes that are included in the genotyping scheme. A sequence type (ST) is assigned to each unique combination of alleles, called an MLST profile. MLST was invented in 1998 and became a de-facto standard taxonomy of bacterial strains, albeit at low resolution. Core genome MLSTAn extension of MLST that analyzes sequence variation across hundreds to thousands of conserved (core) genes, shared by all strains of a species, providing higher resolution typing for genomic epidemiology and evolutionary studies. cgMLST schemes typically comprise 2000 to 4000 genes, depending on the genome size and genetic variation (in terms of presence/absence of genes) within bacterial species. A core genome sequence type (cgST) can be assigned to unique cgMLST profiles, i.e., a unique combination of cgMLST allelic numbers. Whole Genome Sequencing (WGS)A method that determines the complete DNA sequence of an organisms genome in a single process, providing comprehensive information for comparative genetic analyses based on cgMLST or other analytic methods. Single Nucleotide Polymorphisms (SNPs)Variations at a single base position in the DNA sequence among individuals isolates, strains or species, used as genetic markers for studying for example, evolutionary relationships or strain identity. Average nucleotide identity (ANI)A measure of genomic similarity between two organisms, calculated as the average percentage of identical nucleotides in orthologous genomic regions; commonly used to assess species-level relatedness in prokaryotes. TaxonomyHere, we apply the word taxonomy to bacterial strains as a system of classifying, naming and identifying strains based on shared genetic characteristics as defined by e.g., cgMLST.

microbiology↗

A high-resolution gene expression atlas of the medial and lateral domains of the gynoecium of Arabidopsis

Angiosperms are characterized by the formation of flowers, and in their inner floral whorl, one or various gynoecia are produced. These female reproductive structures are responsible for fruit and seed production, thus ensuring the reproductive competence of angiosperms. In Arabidopsis thaliana, the gynoecium is composed of two fused carpels with different tissues that need to develop and differentiate to consolidate a mature gynoecium and thus the reproductive competence of Arabidopsis. For these reasons, they have become the object of study for floral and fruit development. However, due to the complexity of the gynoecium, specific spatio-temporal tissues expression patterns are still scarce. In this study, we used precise laser-assisted microdissection and high-throughput RNA sequencing to describe the transcriptional profiles of the medial and lateral domain tissues of the Arabidopsis gynoecium. We provide evidence that the method used is reliable and that, in addition to corroborating gene expression patterns of previously reported regulators of these tissues, we found genes whose expression dynamics point to being involved in cytokinin and auxin homeostasis and in cell cycle progression. Furthermore, based on differential gene expression analyses, we functionally characterized several genes and found that they are involved in gynoecium development. This new resource is available via the Arabidopsis eFP browser and will serve the community in future studies on developmental and reproductive biology.

plant biology↗