bioRxiv ScienceSearch

Biology subjects

Golubchik, T.

Publications and source records attributed to Golubchik, T..

5 recordsLinked to original sources

A comprehensive genomics solution for HIV surveillance and clinical monitoring in a global health setting

High-throughput viral genetic sequencing is needed to monitor the spread of drug resistance, direct optimal antiretroviral regimes, and to identify transmission dynamics in generalised HIV epidemics. Public health efforts to sequence HIV genomes at scale face three major technical challenges: (i) minimising assay cost and protocol complexity, (ii) maximising sensitivity, and (iii) recovering accurate and unbiased sequences of both the genome consensus and the within-host viral diversity. Here we present a novel, high-throughput, virus-enriched sequencing method and computational pipeline tailored specifically to HIV (veSEQ-HIV), which addresses all three technical challenges, and can be used directly on leftover blood drawn for routine CD4 testing. We demonstrate its performance on 1,620 plasma samples collected from consenting individuals attending 10 large urban clinics in Zambia, partners of HPTN 071 (PopART). We show that veSEQ-HIV consistently recovers complete HIV genomes from the majority of samples of different subtypes, and is also quantitative: the number of HIV reads per sample obtained by veSEQ-HIV estimates viral load without the need for additional testing. Both quantitativity and sensitivity were assessed on a subset of 126 samples with clinically measured viral loads, and with standardized quantification controls (VL 100 - 5,000,000 RNA copies/ml). Complete HIV genomes were recovered from 93% (85/91) of samples when viral load was over 1,000 copies per ml. The quantitative nature of the assay implies that variant frequencies estimated with veSEQ-HIV are representative of true variant frequencies in the sample. Detection of minority variants can be exploited for epidemiological analysis of transmission and drug resistance, and we show how the information contained in individual reads of a veSEQ-HIV sample can be used to detect linkage between multiple mutations associated with resistance to antiretroviral therapy. Less than 2% of reads obtained by veSEQ-HIV were identified as in silico contamination events using updates to the phyloscanner software (phyloscanner clean) that we show to be 95% sensitive and 99% specific at decontaminating NGS data. The cost of the assay -- approximately 45 USD per sample -- compares favourably with existing VL and HIV genotyping tests, and provides the additional value of viral load quantification and inference of drug resistance with a single test. veSEQ-HIV is well suited to large public health efforts and is being applied to all [~]9000 samples collected for the HPTN 071-2 (PopART Phylogenetics) study.

genomics

Human Herpes Virus 6 (HHV-6) - Pathogen or Passenger? A pilot study of clinical laboratory data and next generation sequencing

ABSTRACT\n\nBackgroundHuman herpes virus 6 (HHV-6) is a ubiquitous organism that can cause a variety of clinical syndromes ranging from short-lived rash and fever through to life-threatening encephalitis.\n\nObjectivesWe set out to generate observational data regarding the epidemiology of HHV-6 infection in clinical samples from a UK teaching hospital and to compare different diagnostic approaches.\n\nStudy designFirst, we scrutinized HHV-6 detection in samples submitted to our hospital laboratory through routine diagnostic pathways. Second, we undertook a pilot study using Illumina next generation sequencing (NGS) to determine the frequency of HHV-6 in CSF and respiratory samples that were initially submitted to the laboratory for other diagnostic tests.\n\nResultsOf 72 samples tested for HHV-6 by PCR at the request of a clinician, 24 (33%) were positive for HHV-6. The majority of these patients were under the care of the haematology team (30/41, 73%), and there was a borderline association between HHV-6 detection and both Graft versus Host Disease (GvHD) and Central nervous system (CNS) disease (p=0.05 in each case). We confirmed detection of HHV-6 DNA using NGS in 4/20 (20%) CSF and respiratory samples.\n\nConclusionsHHV-6 is common in clinical samples submitted from a high-risk haematology population, and enhanced screening of this group should be considered. NGS can be used to identify HHV-6 from a complex microbiomee, but further controls are required to define the sensitivity and specificity, and to correlate these results with clinical disease. Our results underpin ongoing efforts to develop NGS technology for viral diagnostics.

microbiology

PHYLOSCANNER: Analysing Within- and Between-Host Pathogen Genetic Diversity to Identify Transmission, Multiple Infection, Recombination and Contamination

A central feature of pathogen genomics is that different infectious particles (virions, bacterial cells, etc.) within an infected individual may be genetically distinct, with patterns of relatedness amongst infectious particles being the result of both within-host evolution and transmission from one host to the next. Here we present a new software tool, phyloscanner, which analyses pathogen diversity from multiple infected hosts. phyloscanner provides unprecedented resolution into the transmission process, allowing inference of the direction of transmission from sequence data alone. Multiply infected individuals are also identified, as they harbour subpopulations of infectious particles that are not connected by within-host evolution, except where recombinant types emerge. Low-level contamination is flagged and removed. We illustrate phyloscanner on both viral and bacterial pathogens, namely HIV-1 sequenced on Illumina and Roche 454 platforms, HCV sequenced with the Oxford Nanopore MinION platform, and Streptococcus pneumoniae with sequences from multiple colonies per individual. phyloscanner is available from https://github.com/BDI-pathogens/phyloscanner.

evolutionary biology

Severe infections emerge from the microbiome by adaptive evolution

Bacteria responsible for the greatest global mortality colonize the human microbiome far more frequently than they cause severe infections. Whether mutation and selection within the microbiome accompany infection is unknown. We investigated de novo mutation in 1163 Staphylococcus aureus genomes from 105 infected patients with nose-colonization. We report that 72% of infections emerged from the microbiome, with infecting and nose-colonizing bacteria showing parallel adaptive differences. We found 2.8-to-3.6-fold enrichments of protein-altering variants in genes responding to rsp, which regulates surface antigens and toxicity; agr, which regulates quorum-sensing, toxicity and abscess formation; and host-derived antimicrobial peptides. Adaptive mutations in pathogenesis-associated genes were 3.1-fold enriched in infecting but not nose-colonizing bacteria. None of these signatures were observed in healthy carriers nor at the species-level, suggesting disease-associated, short-term, within-host selection pressures. Our results show that infection, like a cancer of the microbiome, emerges through spontaneous adaptive evolution, raising new possibilities for diagnosis and treatment.\n\nOne Sentence SummaryLife-threatening S. aureus infections emerge from nose microbiome bacteria in association with repeatable adaptive evolution.

genomics

Easy and Accurate Reconstruction of Whole HIV Genomes from Short-Read Sequence Data

Next-generation sequencing has yet to be widely adopted for HIV. The difficulty of accurately reconstructing the consensus sequence of a quasispecies from reads (short fragments of DNA) in the presence of rapid between- and within-host evolution may have presented a barrier. In particular, mapping (aligning) reads to a reference sequence leads to biased loss of information; this bias can distort epidemiological and evolutionary conclusions. De novo assembly avoids this bias by effectively aligning the reads to themselves, producing a set of sequences called contigs. However contigs provide only a partial summary of the reads, misassembly may result in their having an incorrect structure, and no information is available at parts of the genome where contigs could not be assembled. To address these problems we developed the tool shiver to preprocess reads for quality and contamination, then map them to a reference tailored to the sample using corrected contigs supplemented with existing reference sequences. Run with two commands per sample, it can easily be used for large heterogeneous data sets. We use shiver to reconstruct the consensus sequence and minority variant information from paired-end short-read data produced with the Illumina platform, for 65 existing publicly available samples and 50 new samples. We show the systematic superiority of mapping to shivers constructed reference over mapping the same reads to the standard reference HXB2: an average of 29 bases per sample are called differently, of which 98.5% are supported by higher coverage. We also provide a practical guide to working with imperfect contigs.

bioinformatics