bioRxiv ScienceSearch

Biology subjects

Fraser, C.

Publications and source records attributed to Fraser, C..

10 recordsLinked to original sources

Prediction of post-vaccine population structure of Streptococcus pneumoniae using accessory gene frequencies

Predicting how pathogen populations will change over time is challenging. Such has been the case with Streptococcus pneumoniae, an important human pathogen, and the pneumococcal conjugate vaccines (PCVs), which target only a fraction of the strains in the population. Here, we use the frequencies of accessory genes to predict changes in the pneumococcal population after vaccination, hypothesizing that these frequencies reflect negative frequency-dependent selection (NFDS) on the gene products. We find that the standardized predicted fitness of a strain estimated by an NFDS-based model at the time the vaccine is introduced enables to predict whether the strain increases or decreases in prevalence following vaccination. Further, we are able to forecast the equilibrium post-vaccine population composition and assess the invasion capacity of emerging lineages. Overall, we provide a method for predicting the impact of an intervention on pneumococcal populations with potential application to other bacterial pathogens in which NFDS is a driving force.

evolutionary biology

Link between the numbers of particles and variants founding new HIV-1 infections depends on the timing of transmission

Understanding which HIV-1 variants are most likely to be transmitted is important for vaccine design and predicting virus evolution. Since most infections are founded by single variants, it has been suggested that selection at transmission has a key role in governing which variants are transmitted. We show that the composition of the viral population within the donor at the time of transmission is also important. To support this argument, we developed a probabilistic model describing HIV-1 transmission in an untreated population, and parameterised the model using both within-host next generation sequencing data and population-level epidemiological data on heterosexual transmission. The most basic HIV-1 transmission models cannot explain simultaneously the low probability of transmission and the non-negligible proportion of infections founded by multiple variants. In our model, transmission can only occur when environmental conditions are appropriate (e.g. abrasions are present in the genital tract of the potential recipient), allowing these observations to be reconciled. As well as reproducing features of transmission in real populations, our model demonstrates that, contrary to expectation, there is not a simple link between the number of viral variants and the number of viral particles founding each new infection. These quantities depend on the timing of transmission, and infections can be founded with small numbers of variants yet large numbers of particles. Including selection, or a bias towards early transmission (e.g. due to treatment) acts to enhance this conclusion. In addition, we find that infections initiated by multiple variants are most likely to have derived from donors with intermediate set-point viral loads, and not from individuals with high set-point viral loads as might be expected. We therefore emphasise the importance of considering viral diversity in donors, and the timings of transmissions, when trying to discern the complex factors governing single or multiple variant transmission.

epidemiology

A comprehensive genomics solution for HIV surveillance and clinical monitoring in a global health setting

High-throughput viral genetic sequencing is needed to monitor the spread of drug resistance, direct optimal antiretroviral regimes, and to identify transmission dynamics in generalised HIV epidemics. Public health efforts to sequence HIV genomes at scale face three major technical challenges: (i) minimising assay cost and protocol complexity, (ii) maximising sensitivity, and (iii) recovering accurate and unbiased sequences of both the genome consensus and the within-host viral diversity. Here we present a novel, high-throughput, virus-enriched sequencing method and computational pipeline tailored specifically to HIV (veSEQ-HIV), which addresses all three technical challenges, and can be used directly on leftover blood drawn for routine CD4 testing. We demonstrate its performance on 1,620 plasma samples collected from consenting individuals attending 10 large urban clinics in Zambia, partners of HPTN 071 (PopART). We show that veSEQ-HIV consistently recovers complete HIV genomes from the majority of samples of different subtypes, and is also quantitative: the number of HIV reads per sample obtained by veSEQ-HIV estimates viral load without the need for additional testing. Both quantitativity and sensitivity were assessed on a subset of 126 samples with clinically measured viral loads, and with standardized quantification controls (VL 100 - 5,000,000 RNA copies/ml). Complete HIV genomes were recovered from 93% (85/91) of samples when viral load was over 1,000 copies per ml. The quantitative nature of the assay implies that variant frequencies estimated with veSEQ-HIV are representative of true variant frequencies in the sample. Detection of minority variants can be exploited for epidemiological analysis of transmission and drug resistance, and we show how the information contained in individual reads of a veSEQ-HIV sample can be used to detect linkage between multiple mutations associated with resistance to antiretroviral therapy. Less than 2% of reads obtained by veSEQ-HIV were identified as in silico contamination events using updates to the phyloscanner software (phyloscanner clean) that we show to be 95% sensitive and 99% specific at decontaminating NGS data. The cost of the assay -- approximately 45 USD per sample -- compares favourably with existing VL and HIV genotyping tests, and provides the additional value of viral load quantification and inference of drug resistance with a single test. veSEQ-HIV is well suited to large public health efforts and is being applied to all [~]9000 samples collected for the HPTN 071-2 (PopART Phylogenetics) study.

genomics

Mixing patterns of HIV transmission among men who have sex with men in the United Kingdom

BackgroundNear 60% of new HIV infections in the United Kingdom are estimated to occur in men who have sex with men (MSM). Patterns of mixing between different risk groups of MSM have been suggested to spread the HIV epidemics through age-disassortative partnerships and to contribute to ethnic disparities in infection rates. Understanding these mixing patterns in transmission can help to determine which groups are at a greater risk and guide prevention.\n\nMethodsWe analyzed combined epidemiologic data and viral sequences from MSM diagnosed with HIV as of mid-2015 at the national level. We applied a phylodynamic source attribution model to infer patterns of transmission between groups of patients by age, ethnicity and region.\n\nResultsFrom pair probabilities of transmission between 19 847 MSM patients, we found that potential transmitters of HIV subtype B were on average 5 months older than recipients. We also found a moderate overall assortativity of transmission by ethnic group and a stronger assortativity by region.\n\nConclusionsOur findings suggest that there is only a modest net flow of transmissions from older to young MSM in subtype B epidemics and that young MSM, both for Black or White groups, are more likely to be infected by one another than expected in a sexual network with random mixing.

epidemiology

Mechanisms that maintain coexistence of antibiotic sensitivity and resistance also promote high frequencies of multidrug resistance

Resistance against different antibiotics appears on the same bacterial strains more often than expected by chance, leading to high frequencies of multidrug resistance. There are multiple explanations for this observation, but these tend to be specific to subsets of antibiotics and/or bacterial species, whereas the trend is pervasive. Here, we consider the question in terms of strain ecology: explaining why resistance to different antibiotics is often seen on the same strain requires an understanding of the competition between strains with different resistance profiles. This work builds on models originally proposed to explain another aspect of strain competition: the stable coexistence of antibiotic sensitivity and resistance observed in a number of bacterial species. We first demonstrate a partial structural similarity in these models of coexistence. We then generalise this unified underlying model to multidrug resistance and show that models with this structure predict high levels of association between resistance to different drugs and high multidrug resistance frequencies. We test predictions from this model in six bacterial datasets and find them to be qualitatively consistent with observed trends. The higher than expected frequencies of multidrug resistance are often interpreted as evidence that these strains are out-competing strains with lower resistance multiplicity. Our work provides an alternative explanation that is compatible with long-term stability in resistance frequencies.\n\nAuthor summaryAntibiotic resistance is a serious public health concern, yet the ecology and evolution of drug resistance are not fully understood. This impacts our ability to design effective interventions to combat resistance. From a public health point of view, multidrug resistance is particularly problematic because resistance to different antibiotics is often seen on the same bacterial strains, which leads to high frequencies of multidrug resistance and limits treatment options. This work seeks to explain this trend in terms of strain ecology and the competition between strains with different resistance profiles. Building on recent work exploring why resistant bacteria are not out-competing sensitive bacteria, we show that models originally proposed to explain this observation also predict high multidrug resistance frequencies. These models are therefore a unifying explanation for two pervasive trends in resistance dynamics. In terms of public health, the implication of our results is that new resistances are likeliest to be found on already multidrug resistant strains and that changing patterns of prescription may not be enough to combat multidrug resistance.

evolutionary biology

PHYLOSCANNER: Analysing Within- and Between-Host Pathogen Genetic Diversity to Identify Transmission, Multiple Infection, Recombination and Contamination

A central feature of pathogen genomics is that different infectious particles (virions, bacterial cells, etc.) within an infected individual may be genetically distinct, with patterns of relatedness amongst infectious particles being the result of both within-host evolution and transmission from one host to the next. Here we present a new software tool, phyloscanner, which analyses pathogen diversity from multiple infected hosts. phyloscanner provides unprecedented resolution into the transmission process, allowing inference of the direction of transmission from sequence data alone. Multiply infected individuals are also identified, as they harbour subpopulations of infectious particles that are not connected by within-host evolution, except where recombinant types emerge. Low-level contamination is flagged and removed. We illustrate phyloscanner on both viral and bacterial pathogens, namely HIV-1 sequenced on Illumina and Roche 454 platforms, HCV sequenced with the Oxford Nanopore MinION platform, and Streptococcus pneumoniae with sequences from multiple colonies per individual. phyloscanner is available from https://github.com/BDI-pathogens/phyloscanner.

evolutionary biology

Coalescent models for populations with time-varying population sizes and arbitrary offspring distributions

The coalescent has been used to infer from gene genealogies the population dynamics of biological systems, such as the prevalence of an infectious disease. The offspring distribution affects the relationship between population dynamics and the genealogy, and for infectious diseases, the offspring distribution is often highly overdispersed. Here, we provide a general formula for the coalescent rate for populations with time-varying sizes and any offspring distribution. The formula is valid in the same large population limit as Kingmans original derivation. By relating our derivation to existing formulations of the coalescent, we show that differences in the coalescent rate derived for many population models may be explained by differences in the offspring distribution. The coalescent derivations presented here could be used to quantify the overdispersion in the offspring distribution of infectious diseases, which is useful for accurate modelling disease outbreaks.

evolutionary biology

Host population structure and treatment frequency maintain balancing selection on drug resistance

It is a truism that antimicrobial drugs select for resistance, but explaining pathogen- and population-specific variation in patterns of resistance remains an open problem. Like other common commensals, Streptococcus pneumoniae has demonstrated persistent coexistence of drug-sensitive and drug-resistant strains. Theoretically, this outcome is unlikely. We modeled the dynamics of competing strains of S. pneumoniae to investigate the impact of transmission dynamics and treatment-induced selective pressures on the probability of stable coexistence. We find that the outcome of competition is extremely sensitive to structure in the host population, although coexistence can arise from age-assortative transmission models with age-varying rates of antibiotic use. Moreover, we find that the selective pressure from antibiotics arises not so much from the rate of antibiotic use per se but from the frequency of treatment: frequent antibiotic therapy disproportionately impacts the fitness of sensitive strains. This same phenomenon explains why serotypes with longer durations of carriage tend to be more resistant. These dynamics may apply to other potentially pathogenic, microbial commensals and highlight how population structure, which is often omitted from models, can have a large impact.

evolutionary biology

Frequent recombination of pneumococcal capsule highlights future risks of emergence of novel serotypes.

Capsular diversity of Streptococcus pneumoniae constitutes a major obstacle in eliminating the pneumococcal disease. Such diversity is genetically encoded by almost 100 variants of the capsule polysaccharide locus (cps). However, the evolutionary dynamics of the capsule - the target of the currently used vaccines - remains not fully understood. Here, using genetic data from 4,469 bacterial isolates, we found cps to be an evolutionary hotspot with elevated substitution and recombination rates. These rates were a consequence of altered selection at this locus, supporting the hypothesis that the capsule has an increased potential to generate novel diversity compared to the rest of the genome. Analysis of twelve serogroups revealed their complex evolutionary history, which was principally driven by recombination with other serogroups and other streptococci. We observed significant variation in recombination rates between different serogroups. This variation could only be partially explained by the lineage-specific recombination rate, the remaining factors being likely driven by serogroup-specific ecology and epidemiology. Finally, we discovered two previously unobserved mosaic serotypes in the densely sampled collection from Mae La, Thailand, here termed 10X and 21X. Our results thus emphasise the strong adaptive potential of the bacterium by its ability to generate novel serotypes by recombination.

evolutionary biology

Easy and Accurate Reconstruction of Whole HIV Genomes from Short-Read Sequence Data

Next-generation sequencing has yet to be widely adopted for HIV. The difficulty of accurately reconstructing the consensus sequence of a quasispecies from reads (short fragments of DNA) in the presence of rapid between- and within-host evolution may have presented a barrier. In particular, mapping (aligning) reads to a reference sequence leads to biased loss of information; this bias can distort epidemiological and evolutionary conclusions. De novo assembly avoids this bias by effectively aligning the reads to themselves, producing a set of sequences called contigs. However contigs provide only a partial summary of the reads, misassembly may result in their having an incorrect structure, and no information is available at parts of the genome where contigs could not be assembled. To address these problems we developed the tool shiver to preprocess reads for quality and contamination, then map them to a reference tailored to the sample using corrected contigs supplemented with existing reference sequences. Run with two commands per sample, it can easily be used for large heterogeneous data sets. We use shiver to reconstruct the consensus sequence and minority variant information from paired-end short-read data produced with the Illumina platform, for 65 existing publicly available samples and 50 new samples. We show the systematic superiority of mapping to shivers constructed reference over mapping the same reads to the standard reference HXB2: an average of 29 bases per sample are called differently, of which 98.5% are supported by higher coverage. We also provide a practical guide to working with imperfect contigs.

bioinformatics