bioRxiv Science⌕ Search

Biology subjects

Krisna, M. A.

Publications and source records attributed to Krisna, M. A..

2 recordsLinked to original sources

Neisseria gonorrhoeae LIN codes: a Robust, Multi-Resolution Lineage Nomenclature

Investigation of the bacterial pathogen Neisseria gonorrhoeae is complicated by extensive horizontal gene transfer: a process which disrupts phylogenetic signals and impedes our understanding of population structure. The ability to consistently identify N. gonorrhoeae lineages is important for surveillance of this increasingly antimicrobial resistant organism, facilitating efficient communication regarding its epidemiology; however, conventional typing systems fail to reflect N. gonorrhoeae strain taxonomy in a reliable and stable manner. Here, a N. gonorrhoeae genomic lineage nomenclature, based on the barcoding system of Life Identification Number (LIN) codes, was developed using a refined 1430 core gene MLST (cgMLST). This hierarchical LIN code nomenclature conveys lineage information at multiple levels of resolution within one code, enabling it to provide immediate context to an isolates ancestry, and to relate to familiar, previously used typing schemes such as Ng cgMLST v1, 7-locus MLST, or NG-STAR clonal complex (CC). Clustering with LIN codes accurately reflects gonococcal diversity and population structure, providing insight into associations between genotype and phenotype for traits such as antibiotic resistance. These codes are automatically assigned and publicly accessible via the pubmlst.org/organisms/neisseria-spp database.

genomics↗

Development and Implementation of a Core Genome Multilocus Sequence Typing (cgMLST) scheme for Haemophilus influenzae

2.Haemophilus influenzae is part of the human nasopharyngeal microbiota and a pathogen causing invasive disease. The extensive genetic diversity observed in H. influenzae necessitates discriminatory analytical approaches to evaluate its population structure. This study developed a core genome MLST (cgMLST) scheme for H. influenzae using pangenome analysis tools and validated the cgMLST scheme using datasets consisting of complete reference genomes (N=14) and high-quality draft H. influenzae genomes (N=2,297). The draft genome dataset was divided into a development (N=921) and a validation dataset (N=1,376). The development dataset was used to identify potential core genes with the validation dataset used to refine the final core gene list to ensure the reliability of the proposed cgMLST scheme. Functional classifications were made for all resulting core genes. Phylogenetic analyses were performed using both allelic profiles and nucleotide sequence alignments of the core genome to test congruence, as assessed by Spearmans correlation and Ordinary Least Square linear regression tests. Preliminary analyses using the development dataset identified 1,067 core genes, which were refined to 1,037 with the validation dataset. More than 70% of core genes were predicted to encode proteins essential for metabolism or genetic information processing. Phylogenetic and statistical analyses indicated that the core genome allelic profile accurately represented phylogenetic relatedness among the isolates (R2 = 0.945). We used this cgMLST scheme to define a high-resolution population structure for H. influenzae, which enhances the genomic analysis of this clinically relevant human pathogen. 3. Impact statementDiscriminating H. influenzae variants and evaluating population structure has been challenging and largely unstandardised. To address this, we have developed a cgMLST scheme for H. influenzae. Since an accurate typing approach relies on precise reflection of the underlying population structure, we explored various methods to define the scheme. The core genes included in this scheme were predicted to encode functions in essential biological pathways, such as metabolism and genetic information processing, and could be reliably assembled from short-read sequence data. Single-linkage clustering, based on core genome allelic profiles, showed high congruence to genealogy reconstructed by Maximum-Likelihood (ML) methods from the core genome nucleotide alignment. The cgMLST scheme v1 enables rapid and accurate depiction of high-resolution H. influenzae population structure, and making this scheme accessible via the PubMLST database, ensures that microbiology reference laboratories and public health authorities worldwide can use it for genomic surveillance. 4. Data summaryThe H. influenzae cgMLST scheme is accessible via https://pubmlst.org/organisms/haemophilus-influenzae. The list of isolate IDs available publicly from pubmlst.org is provided in Supplementary File 1. The pipeline for cgMLST scheme development and validation is published at https://www.protocols.io/private/EF6DB7FE429311EEB8630A58A9FEAC02. All in-house R and Python scripts for data processing and analysis are available from https://gitfront.io/r/user-4399403/ZHt8DArALHcY/cgmlst-hinf/.

genomics↗