bioRxiv ScienceSearch

Biology subjects

Darling, A. E.

Publications and source records attributed to Darling, A. E..

7 recordsLinked to original sources

Whole genome analysis of ExPEC ST73 from a single hospital over a 2-year period identified different circulating clonal groups

ST73 has emerged as one of the most frequently isolated extraintestinal pathogenic E. coli (ExPEC). To examine the localised diversity of ST73 clonal groups including their mobile genetic elements profile, we sequenced the genomes of 16 multiple drug-resistant ST73 isolates from patients with urinary tract infection from a single hospital in Sydney, Australia between 2009 and 2011. Genome sequences were used to generate a SNP-based phylogenetic tree to determine the relationship of these isolates in a global context with ST73 sequences (n=210) from public databases. There was no evidence of a dominant outbreak strain of ST73 in patients from this hospital, rather we identified at least eight separate groups, several of which reoccur, over a two-year period. The inferred phylogeny of all ST73 strains (n=226) including the ST73 Clone D i2 reference genome shows high bootstrap support and clusters into four major groups which correlate with serotype. The Sydney ST73 strains carry a wide variety of virulence-associated genes but the presence of iss, pic and several iron acquisition operons was notable.\n\nImpactST73 is a major clonal lineage of ExPEC that causes urinary tract infections often with uroseptic sequelae but has not garnered substantial scientific interest as the globally disseminated ST131. Isolation of multiple antimicrobial resistant variants of ExPEC ST73 have increased in frequency, but little is known about the carriage of class 1 integrons in this sequence type and the plasmids that are likely to mobilise them. This pilot study examines the ST73 isolates within a single hospital in Sydney Australia and provides the first large-scale core-genome phylogenetic analysis of ST73 utilizing public sequence read datasets. We used this analysis to identify at least 8 sub-groups of ST73 within this single hospital. Mobile genetic elements associated with antibiotic resistance were less diverse and only three class 1 integron structures were identified, all sharing the same basic structure suggesting that the acquisition of drug resistance is a recent event. Genomic epidemiological studies are needed to further characterise established and emerging clonal populations of multiple drug resistant ExPEC to identify sources and aid outbreak investigations.

genomics

bin3C : Exploiting Hi-C sequencing data to accurately resolve metagenome-assembled genomes (MAGs)

Most microbes inhabiting the planet cannot be easily grown in the lab. Metagenomic techniques provide a means to study these organisms, and recent advances in the field have enabled the resolution of individual genomes from metagenomes, so-called Metagenome Assembled Genomes (MAGs). In addition to expanding the catalog of known microbial diversity, the systematic retrieval of MAGs stands as a tenable divide and conquer reduction of metagenome analysis to the simpler problem of single genome analysis. Many leading approaches to MAG retrieval depend upon time-series or transect data, whose effectiveness is a function of community complexity, target abundance and depth of sequencing. Without the need for time-series data, promising alternative methods are based upon the high-throughput sequencing technique called Hi-C.\n\nThe Hi-C technique produces read-pairs which capture in-vivo DNA-DNA proximity interactions (contacts). The physical structure of the community modulates the signal derived from these interactions and a hierarchy of interaction rates exists ([i]ntra-chromosomal > Inter-chromosomal > Inter-cellular).\n\nWe describe an unsupervised method that exploits the hierarchical nature of Hi-C interaction rates to resolve MAGs from a single time-point. As a quantitative demonstration, next, we validate the method against the ground truth of a simulated human faecal microbiome. Lastly, we directly compare our method against a recently announced proprietary service ProxiMeta, which also performs MAG retrieval using Hi-C data.\n\nbin3C has been implemented as a simple open-source pipeline and makes use of the unsupervised community detection algorithm Infomap (https://github.com/cerebis/bin3C).

bioinformatics

CAMISIM: Simulating metagenomes and microbial communities

Shotgun metagenome data sets of microbial communities are highly diverse, not only due to the natural variation of the underlying biological systems, but also due to differences in laboratory protocols, replicate numbers, and sequencing technologies. Accordingly, to effectively assess the performance of metagenomic analysis software, a wide range of benchmark data sets are required. Here, we describe the CAMISIM microbial community and metagenome simulator. The software can model different microbial abundance profiles, multi-sample time series and differential abundance studies, includes real and simulated strain-level diversity, and generates second and third generation sequencing data from taxonomic profiles or de novo. Gold standards are created for sequence assembly, genome binning, taxonomic binning, and taxonomic profiling. CAMSIM generated the benchmark data sets of the first CAMI challenge. For two simulated multi-sample data sets of the human and mouse gut microbiomes we observed high functional congruence to the real data. As further applications, we investigated the effect of varying evolutionary genome divergence, sequencing depth, and read error profiles on two popular metagenome assemblers, MEGAHIT and metaSPAdes, on several thousand small data sets generated with CAMISIM. CAMISIM can simulate a wide variety of microbial communities and metagenome data sets together with truth standards for method evaluation. All data sets and the software are freely available at: https://github.com/CAMI-challenge/CAMISIM

bioinformatics

Porcine commensal Escherichia coli: A reservoir for class 1 integrons associated with IS26

Porcine faecal waste is a serious environmental pollutant. Carriage of antimicrobial resistance and virulence-associated genes (VAGs) and the zoonotic potential of commensal Escherichia coli from swine is largely unknown. Furthermore, little is known about the role of commensal E. coli as contributors to the mobilisation of antimicrobial resistance genes between food animals and the environment. Here, we report whole genome sequence analysis of 141 E. coli from the faeces of healthy pigs. Most strains belonged to phylogroups A and B1 and carried i) a class 1 integron; ii) VAGs linked with extraintestinal infection in humans; iii) antimicrobial resistance genes blaTEM, aphAl, cmlA, strAB, tet(A)A, dfrA12, dfrA5, sul1, sul2, sul3; iv) IS26; and v) heavy metal resistance genes (merA, cusA, terA). Carriage of the sulphonamide resistance gene sul3 was notable in this study. The 141 strains belonged to 42 multilocus sequence types, but clonal complex 10 featured prominently. Structurally diverse class 1 integrons that were frequently associated with IS26 carried unique genetic features that were also identified in extraintestinal pathogenic E. coli (ExPEC) from humans. This study provides the first detailed genomic analysis and point of reference for commensal E. coli of porcine origin, facilitating tracking of specific lineages and the mobile resistance genes they carry.\n\nConflict of Interest StatementNone to declare.

genomics

Effective Online Bayesian Phylogenetics Via Sequential Monte Carlo With Guided Proposals

AO_SCPLOWBSTRACTC_SCPLOWModern infectious disease outbreak surveillance produces continuous streams of sequence data which require phylogenetic analysis as data arrives. Current software packages for Bayesian phy-logenetic inference are unable to quickly incorporate new sequences as they become available, making them less useful for dynamically unfolding evolutionary stories. This limitation can be addressed by applying a class of Bayesian statistical inference algorithms called sequential Monte Carlo (SMC) to conduct online inference, wherein new data can be continuously incorporated to update the estimate of the posterior probability distribution. In this paper we describe and evaluate several different online phylogenetic sequential Monte Carlo (OPSMC) algorithms. We show that proposing new phylogenies with a density similar to the Bayesian prior suffers from poor performance, and we develop guided proposals that better match the proposal density to the posterior. Furthermore, we show that the simplest guided proposals can exhibit pathological behavior in some situations, leading to poor results, and that the situation can be resolved by heating the proposal density. The results demonstrate that relative to the widely-used MCMC-based algorithm implemented in MrBayes, the total time required to compute a series of phylogenetic posteriors as sequences arrive can be significantly reduced by the use of OPSMC, without incurring a significant loss in accuracy.

bioinformatics

Evaluation of ddRADseq for reduced representation metagenome sequencing

Background Who is doing what is the ultimate open question in microbiome study. Shotgun metagenomics is often applied to gain knowledge of functional roles for bacteria in microbial communities, where the data can be used to predict protein encoding genes and enzymatic pathways present in the community, sometimes leading to testable hypotheses for microbial function. We describe a method and basic analysis for a metagenomic adaptation of the double digest restriction site associated DNA sequencing (ddRADseq) protocol for reduced representation metagenome profiling. This technique takes advantage of the sequence specificity of restriction endonucleases to construct an Illumina-compatible sequencing library containing DNA fragments that are between a pair of restriction sites located within close proximity. This results in a reduced sequencing library with coverage breadth that can be tuned by size selection.\n\nResultsWe assessed the performance of the metagenomic ddRADseq approach by applying the method to human stool samples and generating sequence data. We evaluate the extent to which ddRADseq data provides an unbiased reduced representation for microbiome profiling.\n\nConclusionAlthough ddRADseq does introduce some bias in taxonomic representation, the bias is likely to be small relative to DNA extraction bias. ddRADseq appears feasible and could have value as a tool for metagenome-wide association studies.

genomics

Sim3C: Simulation Of HiC And Meta3C Proximity Ligation Sequencing Technologies

BackgroundChromosome conformation capture (3C) and HiC DNA sequencing methods have rapidly advanced our understanding of the spatial organization of genomes and metagenomes. Many variants of these protocols have been developed, each with their own strengths. Currently there is no systematic means for simulating sequence data from this family of sequencing protocols.\n\nFindingsWe describe a computational simulator that, given reference genome sequences and some basic parameters, will simulate HiC sequencing on those sequences. The simulator models the basic spatial structure in genomes that is commonly observed in HiC and 3C datasets, including the distance-decay relationship in proximity ligation, differences in the frequency of interaction within and across chromosomes, and the structure imposed by cells. A means to model the 3D structure of topologically associating domains (TADs) is provided. The simulator also models several sources of error common to 3C and HiC library preparation and sequencing methods, including spurious proximity ligation events and sequencing error.\n\nConclusionsWe have introduced the first comprehensive simulator for 3C and HiC sequencing protocols. We expect the simulator to have use in testing of HiC data analysis algorithms, as well as more general value for experimental design, where questions such as the required depth of sequencing, enzyme choice, and other decisions must be made in advance in order to ensure adequate statistical power to test the relevant hypotheses.

bioinformatics