bioRxiv ScienceSearch

Biology subjects

Carrico, J. A.

Publications and source records attributed to Carrico, J. A..

5 recordsLinked to original sources

The Integrated Rapid Infectious Disease Analysis (IRIDA) Platform

Whole genome sequencing (WGS) is a powerful tool for public health infectious disease investigations owing to its higher resolution, greater efficiency, and cost-effectiveness over traditional genotyping methods. Implementation of WGS in routine public health microbiology laboratories is impeded by a lack of user-friendly automated and semi-automated pipelines, restrictive jurisdictional data sharing policies, and the proliferation of non-interoperable analytical and reporting systems. To address these issues, we developed the Integrated Rapid Infectious Disease Analysis (IRIDA) platform (irida.ca), a user-friendly, decentralized, open-source bioinformatics and analytical web platform to support real-time infectious disease outbreak investigations using WGS data. Instances can be independently installed on local high-performance computing infrastructure, enabling private and secure data management and analyses according to organizational policies and governance. IRIDAs data management capabilities enable secure upload, storage and sharing of all WGS data and metadata. The core platform currently includes pipelines for quality control, assembly, annotation, variant detection, phylogenetic analysis, in silico serotyping, multi-locus sequence typing, and genome distance calculation. Analysis pipeline results can be visualized within the platform through dynamic line lists and integrated phylogenomic clustering for research and discovery, and for enhancing decision-making support and hypothesis generation in epidemiological investigations. Communication and data exchange between instances are provided through customizable access controls. IRIDA complements centralized systems, empowering local analytics and visualizations for genomics-based microbial pathogen investigations. IRIDA is currently transforming the Canadian public health ecosystem and is freely available at https://github.com/phac-nml/irida and www.irida.ca.\n\nImpact StatementWhole genome sequencing (WGS) is revolutionizing infectious disease analysis and surveillance due to its cost effectiveness, utility, and improved analytical power. To date, no \"one-size-fits-all\" genomics platform has been universally adopted, owing to differences in national (and regional) health information systems, data sharing policies, computational infrastructures, lack of interoperability and prohibitive costs. The Integrated Rapid Infectious Disease Analysis (IRIDA) platform is a user-friendly, decentralized, open-source bioinformatics and analytical web platform developed to support real-time infectious disease outbreak investigations using WGS data. IRIDA empowers public health, regulatory and clinical microbiology laboratory personnel to better incorporate WGS technology into routine operations by shielding them from the computational and analytical complexities of big data genomics. IRIDA is now routinely used as part of a validated suite of tools to support outbreak investigations in Canada. While IRIDA was designed to serve the needs of the Canadian public health system, it is generally applicable to any public health and multi-jurisdictional environment. IRIDA enables localized analyses but provides mechanisms and standard outputs to enable data sharing. This approach can help overcome pervasive challenges in real-time global infectious disease surveillance, investigation and control, resulting in faster responses, and ultimately, better public health outcomes.\n\nDATA SUMMARYO_LIData used to generate some of the figures in this manuscript can be found in the NCBI BioProject PRJNA305824.\nC_LI

bioinformatics

Rapid Identification of Stable Clusters in Bacterial Populations Using the Adjusted Wallace Coefficient

Whole-genome sequencing (WGS) of microbial pathogens has become an essential part of modern epidemiological investigations. Although WGS data can be analyzed using a number of different approaches, such as traditional phylogenetic methods, a critical requirement for global systems for pathogen surveillance is the development of approaches for transforming sequence data into WGS-based subtypes, which creates a nomenclature that describes their higher-order relationships to one another. To this end, subtype similarity thresholds are needed to define clusters of subtypes representing lineages of interest. WGS-based subtyping presents a challenge since both the addition of novel genome sequences and small adjustments in similarity thresholds can have a dramatic impact on cluster composition and stability. We present the Neighbourhood Adjusted Wallace Coefficient (nAWC), a method for evaluating cluster stability based on computing cluster concordance between neighbouring similarity thresholds. The nAWC can be used to identify areas in in which distance thresholds produce robust clusters. Using datasets from Salmonella enterica and Campylobacter jejuni, representing strongly and weakly clonal bacterial species respectively, we show that clusters generated using such thresholds are both stable and reflect basic units in their overall population structure. Our results suggest that the nAWC could be useful for defining robust clusters compatible with nomenclatures for global WGS-based surveillance networks, which require stable clusters to be defined that both harness the discriminatory power of WGS data while allowing for long-term tracking of strains of interest.

genomics

Origin, evolution, and distribution of the molecular machinery for biosynthesis of sialylated lipooligosaccharide structures in Campylobacter coli

Campylobacter jejuni and Campylobacter coli are the most common cause of bacterial gastroenteritis worldwide. Additionally, C. jejuni is the most common bacterial etiological agent in the autoimmune Guillain-Barre syndrome (GBS). Ganglioside mimicry by C. jejuni lipooligosaccharide (LOS) is the triggering factor of the disease. LOS-associated genes involved in the synthesis (neuABC) and transfer of sialic acid (sialyltranferases) are essential in C. jejuni to synthesize ganglioside-like LOS. Therefore these genes have been identified as GBS markers. So far, scarce genetic evidence supports C. coli as a GBS causative agent despite being isolated from GBS patients. Here we show that genes putatively involved in sialic acid transfer are widely distributed in the C. coli population. Evidence found herein suggests that a small group of C. coli strains are very likely to express ganglioside mimics, implying that C. coli can potentially trigger GBS. C. coli also presents a larger repertoire of sialyltransferases than C. jejuni and loss of functions of some those LOS-associated genes has happened during adaptation to agriculture niche. Nevertheless, the activity of these sialyltransferases and their role in shaping C. coli population is yet to be explored.

microbiology

GrapeTree: Visualization of core genomic relationships among 100,000 bacterial pathogens

O_LICurrent methods struggle to reconstruct and visualise the genomic relationships of [≥]100,000 bacterial genomes.\nC_LIO_LIGrapeTree facilitates the analyses of allelic profiles from 10,000s of core genomes within a web browser window.\nC_LIO_LIGrapeTree implements a novel minimum spanning tree algorithm to reconstruct genetic relationships despite missing data together with a static \"GrapeTree Layout\" algorithm to render interactive visualisations of large trees.\nC_LIO_LIGrapeTree is a stand-along package for investigating Newick trees plus associated metadata and is also integrated into EnteroBase to facilitate cutting edge navigation of genomic relationships among >160,000 genomes from bacterial pathogens.\nC_LIO_LIThe GrapeTree package was released under the GPL v3.0 Licence.\nC_LI

bioinformatics

chewBBACA: A complete suite for gene-by-gene schema creation and strain identification

Gene-by-gene approaches are becoming increasingly popular in bacterial genomic epidemiology and outbreak detection. However, there is a lack of open-source scalable software for schema definition and allele calling for these methodologies. The chewBBACA suite was designed to assist users in the creation and evaluation of novel whole-genome or core-genome gene-by-gene typing schemas and subsequent allele calling in bacterial strains of interest. The software can run in a laptop or in high performance clusters making it useful for both small laboratories and large reference centers. ChewBBACA is available at https://github.com/B-UMMI/chewBBACA or as a docker image at https://hub.docker.com/r/ummidock/chewbbaca/.\n\nDATA SUMMARYO_LIAssembled genomes used for the tutorial were downloaded from NCBI in August 2016 by selecting those submitted as Streptococcus agalactiae taxon or sub-taxa. All the assemblies have been deposited as a zip file in FigShare (https://figshare.com/s/9cbe1d422805db54cd52), where a file with the original ftp link for each NCBI directory is also available.\nC_LIO_LICode for the chewBBACA suite is available at https://github.com/B-UMMI/chewBBACA while the tutorial example is found at https://github.com/B-UMMI/chewBBACA_tutorial.\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTThe chewBBACA software offers a computational solution for the creation, evaluation and use of whole genome (wg) and core genome (cg) multilocus sequence typing (MLST) schemas. It allows researchers to develop wg/cgMLST schemes for any bacterial species from a set of genomes of interest. The alleles identified by chewBBACA correspond to potential coding sequences, possibly offering insights into the correspondence between the genetic variability identified and phenotypic variability. The software performs allele calling in a matter of seconds to minutes per strain in a laptop but is easily scalable for the analysis of large datasets of hundreds of thousands of strains using multiprocessing options. The chewBBACA software thus provides an efficient and freely available open source solution for gene-by-gene methods. Moreover, the ability to perform these tasks locally is desirable when the submission of raw data to a central repository or web services is hindered by data protection policies or ethical or legal concerns.

bioinformatics