bioRxiv ScienceSearch

bioRxiv · 10.1101/309187

Whole genome analysis of local Kenyan and global sequences unravels the epidemiological and molecular evolutionary dynamics of RSV genotype ON1 strains

Abstract

The respiratory syncytial virus (RSV) group A variant with the 72-nucleotide duplication in the G gene, genotype ON1, was first detected in Kilifi in 2012 and has almost completely replaced previously circulating genotype GA2 strains. This replacement suggests some fitness advantage of ON1 over the GA2 viruses, and might be accompanied by important genomic substitutions in ON1 viruses. Close observation of such a new virus introduction over time provides an opportunity to better understand the transmission and evolutionary dynamics of the pathogen. We have generated and analyzed 184 RSV-A whole genome sequences (WGS) from Kilifi (Kenya) collected between 2011 and 2016, the first ON1 genomes from Africa and the largest collection globally from a single location. Phylogenetic analysis indicates that RSV-A transmission into this coastal Kenya location is characterized by multiple introductions of viral lineages from diverse origins but with varied success in local transmission. We identify signature amino acid substitutions between ON1 and GA2 viruses within genes encoding the surface proteins (G, F), polymerase (L) and matrix M2-1 proteins, some of which were identified as positively selected, and thereby provide an enhanced picture of RSV-A diversity. Furthermore, five of the eleven RSV open reading frames (ORF) (i.e. G, F, L, N and P), analyzed separately, formed distinct phylogenetic clusters for the two genotypes. This might suggest that coding regions outside of the most frequently studied G ORF play a role in the adaptation of RSV to host populations with the alternative possibility that some of the substitutions are nothing more than genetic hitchhikers. Our analysis provides insight into the epidemiological processes that define RSV spread, highlights the genetic substitutions that characterize emerging strains, and demonstrates the utility of large-scale WGS in molecular epidemiological studies.\n\nAuthor summaryRespiratory syncytial virus (RSV) is the leading viral cause of severe pneumonia and bronchiolitis among infants and children globally. No vaccine exists to date. The high genetic variability of this RNA virus, characterized by group (A or B), genotype (within group) and variant (within genotype) replacement in populations, may pose a challenge to effective vaccine design by enabling immune response escape. To date most sequence data exists for the highly variable G gene encoding the RSV attachment protein, and there is little globally-sampled RSV genomic data to provide a fine resolution of the epidemiology and evolutionary dynamics of the pathogen. Here we use long-term RSV surveillance in coastal Kenya to track the introduction, spread and evolution of a new RSV genotype known as ON1 (having a 72-nucleotide duplication in the G gene). We present a set of 184 RSV-A whole genomes, including 176 of RSV ON1 (the first from Africa), describe patterns of local ON1 spread and show genome-wide changes between the two major RSV-A genotypes that may define the pathogens adaptation to the host. These findings have implications for vaccine design and improved understanding of RSV epidemiology and evolution.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Otieno, J. R., Kamau, E. M., Oketch, J. W., Ngoi, J. M., Gichuki, A. M., Binter, S., Otieno, G. P., Ngama, M., Agoti, C. N., Cane, P. A., Kellam, P., Cotten, M., Lemey, P., Nokes, D. J.. 2018-04-27. Whole genome analysis of local Kenyan and global sequences unravels the epidemiological and molecular evolutionary dynamics of RSV genotype ON1 strains. https://doi.org/10.1101/309187

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Translating surveillance data into incidence estimates

Monitoring a population for a disease requires the hosts to be sampled and tested for the pathogen. This results in sampling series from which to estimate the disease incidence, i.e. the proportion of hosts infected. Existing estimation methods assume that disease incidence is not changing between monitoring rounds, resulting in underestimation of the disease incidence. In this paper we develop an incidence estimation model accounting for epidemic growth with monitoring rounds sampling varying incidence. We also show how to accommodate the asymptomatic period characteristic to most diseases. For practical use, we produce an approximation of the model, which is subsequently shown accurate for relevant epidemic and sampling parameters. Both the approximation and the full model are applied to stochastic spatial simulations of epidemics. The results prove their consistency for a very wide range of situations.

epidemiology

The Swiss Primary Ciliary Dyskinesia registry: objectives, methods and first results

Primary Ciliary Dyskinesia (PCD) is a rare hereditary, multi-organ disease caused by defects in ciliary structure and function. It results in a wide range of clinical manifestations, most commonly in the upper and lower airways. Central data collection in national and international registries is essential to studying the epidemiology of rare diseases and filling in gaps in knowledge of diseases such as PCD. For this reason, the Swiss Primary Ciliary Dyskinesia Registry (CH-PCD) was founded in 2013 as a collaborative project between epidemiologists and adult and paediatric pulmonologists.\n\nThe registry records patients of any age, suffering from PCD, who are treated and resident in Switzerland. It collects information from patients identified through physicians, diagnostic facilities, and patient organisations. The registry dataset contains data on diagnostic evaluations, lung function, microbiology and imaging, symptoms, treatments, and hospitalizations.\n\nBy May 2018, CH-PCD has contacted 566 physicians of different specialties and identified 134 patients with PCD. At present this number represents an overall 1 in 63,000 prevalence of people diagnosed with PCD in Switzerland. Prevalence differs by age and region; it is highest in children and adults younger than 30 years, and in Espace Mittelland. The median age of patients in the registry is 25 years (range 5-73), and 49 patients have a definite PCD diagnosis based on recent international guidelines. Data from CH-PCD are contributed to international collaborative studies and the registry facilitates patient identification for nested studies.\n\nCH-PCD has proven to be a valuable research tool that already has highlighted weaknesses in PCD clinical practice in Switzerland. Development of centralised diagnostic and management centres and adherence to international guidelines are needed to improve diagnosis and management--particularly for adult PCD patients.

epidemiology

Perfect Counterfactuals for Epidemic Simulations

Simulation studies are often used to predict the expected impact of control measures in infectious disease outbreaks. Typically, two independent sets of simulations are conducted, one with the intervetnion, and one without, and epidemic sizes (or some related metric) are compared to estimate the effect of the intervention. Since it is possible that controlled epidemics are larger than uncontrolled ones if there is substantial stochastic variation between epidemics, uncertainty intervals from this approach can include a negative effect even for an effective intervention. To more precisely estimate the number of cases an intervention will prevent within a single epidemic, here we develop a single world approach to matching simulations of controlled epidemics to their exact uncontrolled counterfac-tual. Our method borrows concepts from percolation approaches prune out possible epidemic histories and create potential epidemic graph that can be realized to create perfectly matched controlled and uncontrolled epidemics. We present an implementation of this method for a common class of compartmental models, and its application in a simple SIR model. Results illustrate how, at the cost of some computation time, this method substantially narrows confidence intervals and avoids non-sensical inferences.

epidemiology