bioRxiv Science⌕ Search

Biology subjects

Cuddihy, T.

Publications and source records attributed to Cuddihy, T..

3 recordsLinked to original sources

An optimised method for bacterial nucleic acid extraction from positive blood culture broths for whole genome sequencing, resistance phenotype prediction and downstream molecular applications

BackgroundA prerequisite to rapid molecular detection of pathogens causing bloodstream infections is an efficient, cost effective and robust DNA extraction solution. We describe methods for microbial DNA extraction direct from positive blood culture broths, suitable for metagenomic sequencing and the application of machine-learning based tools to predict antimicrobial susceptibility. MethodsProspectively collected culture-positive blood culture broths with Gram-negative bacteria, were directly extracted using various commercially available kits. We compared methods for efficient inhibitor removal, avoidance of DNA shearing or degradation, to achieve DNA of high quality and purity. Bacterial species identified via whole-genome metagenomic sequencing (Illumina, MiniSeq) from blood culture extracts were compared to conventional methods from cultured isolates (Vitek MS). A machine-learning algorithm (AREScloud) was used to predict susceptibility against commercially available antibiotics, compared to susceptibility testing (Vitek 2) and other commercially available rapid diagnostic instruments (Accelerate Pheno and BCID). ResultsA two-kit method using a modified MolYsis Basic kit (for host DNA depletion) and extraction using Qiagen DNeasy UltraClean microbial kits resulted in optimal extractions appropriate for multiple molecular applications, including PCR, short-read and long-read sequencing. DNA extracts from 40 blood culture broths were included. Taxonomic profiling by direct metagenomic sequencing matched species identification by conventional methods in 38/40 (95%) of samples, with two showing agreement to genus level. In two polymicrobial samples, a second organism was missed by sequencing. Whole genome sequencing antimicrobial susceptibility testing (WGS-AST) models were able to accurately infer profiles for 6 common pathogens against 17 antibiotics. Overall categorical agreement (CA) was 95%, with 11% very major errors (VME) and 3.9% major errors (ME). CA for WGS-AST was >95% for 5/6 of the most common pathogens (E. coli, K. pneumoniae, P. mirabilis, P. aeruginosa and C. jejuni) while it was lower for K. oxytoca (66.7%), likely due to the presence of inducible cephalosporinases. Performance of WGS-AST was sub-optimal for uncommon pathogens (e.g. Elizabethkingia) and some combination antibiotic compounds (e.g. ticarcillin-clavulanate). Time to pathogen identification and resistance gene detection was fastest with BCID (1 h from blood culture positivity). Accelerate Pheno provided a rapid MIC result in approximately 8 h. While Illumina based direct metagenomic sequencing did not result in faster turn-around times compared conventional methods, use of real-time nanopore sequencing may allow faster data acquisition. ConclusionsThe application of direct metagenomic sequencing from positive blood culture broths is a feasible approach and solves some of the challenges of sequencing from low-bacterial load samples. Machine-learning based algorithms are also accurate for common pathogen / drug combinations, although additional work is required to optimise algorithms for uncommon species and more complex resistance genotypes, as well as streamlining methods to provide more rapid sequencing results.

microbiology↗

Systematic benchmarking of all-in-one microbial SNP calling pipelines

Clinical and public health microbiology is increasingly utilising whole genome sequencing (WGS) technology and this has lead to the development of a myriad of analysis tools and bioinformatics pipelines. Single nucleotide polymorphism (SNP) analysis is an approach used for strain characterisation and determining isolate relatedness. However, in order to ensure the development of robust methodologies suitable for clinical application of this technology, accurate, reproducible, traceable and benchmarked analysis pipelines are necessary. To date, the approach to benchmarking of these has been largely ad-hoc with new pipelines benchmarked on their own datasets with limited comparisons to previously published pipelines. In this study, Snpdragon, a fast and accurate SNP calling pipeline is introduced. Written in Nextflow, Snpdragon is capable of handling small to very large and incrementally growing datasets. Snpdragon is benchmarked using previously published datasets against six other all-in-one microbial SNP calling pipelines, Lyveset, Lyveset2, Snippy, SPANDx, BactSNP and Nesoni. The effect of dataset choice on performance measures is demonstrated to highlight some of the issues associated with the current available benchmarking approaches. The establishment of an agreed upon gold-standard benchmarking process for microbial variant analysis is becoming increasingly important to aid in its robust application, improve transparency of pipeline performance under different settings and direct future improvements and development. Snpdragon is available at https://github.com/FordeGenomics/SNPdragon. Impact statementWhole-genome sequencing has become increasingly popular in infectious disease diagnostics and surveillance. The resolution provided by single nucleotide polymorphism (SNP) analyses provides the highest level of insight into strain characteristics and relatedness. Numerous approaches to SNP analysis have been developed but with no established gold-standard benchmarking approach, choice of bioinformatics pipeline tends to come down to laboratory or researcher preference. To support the clinical application of this technology, accurate, transparent, auditable, reproducible and benchmarked pipelines are necessary. Therefore, Snpdragon has been developed in Nextflow to allow transparency, auditability and reproducibility and has been benchmarked against six other all-in-one pipelines using a number of previously published benchmarking datasets. The variability of performance measures across different datasets is shown and illustrates the need for a robust, fair and uniform approach to benchmarking. Data SummaryO_LIPreviously sequenced reads for Escherichia coli O25b:H4-ST131 strain EC958 are available in BioProject PRJNA362676. BioSample accession numbers for the three benchmarking isolates are: O_LIEC958: SAMN06245884 C_LIO_LIMS6573: SAMN06245879 C_LIO_LIMS6574: SAMN06245880 C_LI C_LIO_LIAccession numbers for reference genomes against the E. coli O25b:H4-ST131 strain EC958 benchmark are detailed in table 2. C_LIO_LISimulated benchmarking data previously described by Yoshimura et al. is available at http://platanus.bio.titech.ac.jp/bactsnp (1). C_LIO_LISimulated datasets previously described by Bush et al. is available at http://dx.doi.org/10.5287/bodleian:AmNXrjYN8 (2). C_LIO_LIReal sequencing benchmarking datasets previously described by Bush et al. are available at http://dx.doi.org/10.5287/bodleian:nrmv8k5r8 (2). C_LI O_TBL View this table: org.highwire.dtl.DTLVardef@f2446corg.highwire.dtl.DTLVardef@16a299corg.highwire.dtl.DTLVardef@d1c7a5org.highwire.dtl.DTLVardef@8a3cbaorg.highwire.dtl.DTLVardef@198f455_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 2.C_FLOATNO O_TABLECAPTIONReference genomes used in the EC958 benchmarking dataset and the percent identity against the three included samples. C_TABLECAPTION C_TBL

bioinformatics↗

ScrepYard: an online resource for disulfide-stabilised tandem repeat peptides

Receptor avidity through multivalency is a highly sought-after property of ligands. While readily available in nature in the form of bivalent antibodies, this property remains challenging to engineer in synthetic molecules. The discovery of several bivalent venom peptides containing two homologous and independently folded domains (in a tandem repeat arrangement) has provided a unique opportunity to better understand the underpinning design of multivalency in multimeric biomolecules, as well as how naturally occurring multivalent ligands can be identified. In previous work we classified these molecules as a larger class termed secreted cysteine-rich repeat-proteins (SCREPs). Here, we present an online resource; ScrepYard, designed to assist researchers in identification of SCREP sequences of interest and to aid in characterizing this emerging class of biomolecules. Analysis of sequences within the ScrepYard reveals that two-domain tandem repeats constitute the most abundant SCREP domain architecture, while the interdomain "linker" regions connecting the ordered domains are found to be abundant in amino acids with short or polar sidechains and contain an unusually high abundance of proline residues. Finally, we demonstrate the utility of ScrepYard as a virtual screening tool for discovery of putatively multivalent peptides, by using it as a resource to identify a previously uncharacterised serine protease inhibitor and confirm its predicated activity using an enzyme assay.

bioinformatics↗