bioRxiv Science⌕ Search

Biology subjects

Maklin, T.

Publications and source records attributed to Maklin, T..

3 recordsLinked to original sources

Enhanced metagenomics-enabled transmission inference with TRACS

Coexisting strains of the same species within the human microbiota pose a substantial challenge to inferring the host-to-host transmission of both pathogenic and commensal microbes. Here, we present TRACS, a highly accurate algorithm for estimating genetic distances between strains at the level of individual SNPs, which is robust to intra-species diversity within the host. Analysis of well-characterised Faecal Microbiota Transplantation datasets, along with extensive simulations, demonstrates that TRACS substantially outperforms existing strain aware transmission inference methods. We use TRACS to infer transmission networks in patients colonised with multiple strains, including SARS-CoV-2 amplicon sequencing data from UK hospitals, deep population sequencing data of Streptococcus pneumoniae and single-cell genome sequencing data from malaria patients infected with Plasmodium falciparum. Applying TRACS to gut metagenomic samples from a large cohort of 176 mothers and 1,288 infants born in UK hospitals revealed species-specific transmission rates between mothers and their infants. Notably, TRACS identified increased persistence of Bifidobacterium breve in infants, a finding missed by previous analyses due to the presence of multiple strains.

bioinformatics↗

Basic reproduction number for pandemic Escherichia coli clones is comparable to typical pandemic viruses

Extra-intestinal pathogenic Escherichia coli (ExPEC) ubiquitously colonize the human gut and represent clinically the most significant bacterial species causing urinary tract infections and bacteremia in addition to contributing to meningitis in neonates. During the last two decades, new E. coli multi-drug resistant clones such as ST131, particularly its clades C1 and C2, have spread globally, as has their generally less resistant sister clade ST131-A. Phylodynamic coalescent modeling has indicated exponential growth in the populations corresponding to these clades during the early 2000s. However, it remains unknown how their transmission dynamics compare to viral epidemics and pandemics in terms of key epidemiological quantities such as the basic reproduction number (R0). Estimation of R0 for opportunistic pathogenic bacteria poses a difficult challenge compared to viruses causing acute infections, since data on E. coli infections accumulate with a much longer delay, even in the most advanced public health reporting systems. Here, we developed a compartmental model for asymptomatic gut colonization and onward transmission coupled with a stochastic epidemiological observation model for bacteremia and fitted the model to annual Norwegian national E. coli disease surveillance and bacterial population genomics data. Approximate Bayesian Computation leveraged by the ELFI software package was used to infer R0 for the pandemic ST131 clades. The resulting R0 estimates for ST131-A, ST131-C1 and ST131-C2 were 1.47 (1.31-1.60), 1.18 (1.12-1.20) and 1.13 (1.08-1.20), respectively, indicating that the ST131-A transmission potential can be even comparable to pandemic influenza viruses, such as H1N1. The significantly lower transmissibility of ST131-C1 and ST131-C2 suggests that their global dissemination has been aided by antibiotic selection pressure and that they may be more effectively transmitted through healthcare facilities instead of primarily community-driven transmission. In summary our results provide a fundamental advance in understanding the relative transmissibility of these opportunistic pathogens and that it can vary markedly even between very closely related E. coli.

microbiology↗

Themisto: a scalable colored k-mer index for sensitive pseudoalignment against hundreds of thousands of bacterial genomes

MotivationHuge data sets containing whole-genome sequences of bacterial strains are now commonplace and represent a rich and important resource for modern genomic epidemiology and metagenomics. In order to efficiently make use of these data sets, efficient indexing data structures -- that are both scalable and provide rapid query throughput -- are paramount. ResultsHere, we present Themisto, a scalable colored k-mer index designed for large collections of microbial reference genomes, that works for both short and long read data. Themisto indexes 179 thousand Salmonella enterica genomes in 9 hours. The resulting index takes 142 gigabytes. In comparison, the best competing tools Metagraph and Bifrost were only able to index 11 thousand genomes in the same time. In pseudoalignment, these other tools were either an order of magnitude slower than Themisto, or used an order of magnitude more memory. Themisto also offers superior pseudoalignment quality, achieving a higher recall than previous methods on Nanopore read sets. Availability and implementationThemisto is available and documented as a C++ package at https://github.com/algbio/themisto available under the GPLv2 license. Contactjarno.alanko@helsinki.fi Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗