bioRxiv Science⌕ Search

Biology subjects

Herzog, K. S.

Publications and source records attributed to Herzog, K. S..

3 recordsLinked to original sources

Hookworm genomic diversity and population structure from accessible sample types: A validated approach to generate genome-wide polymorphism datasets from individual third-stage larvae

Hookworm infection is a neglected tropical disease affecting hundreds of thousands of people annually in the tropics and sub-tropics. Population genomic approaches have the potential to improve our understanding of hookworm infection dynamics and control efficacy. Here, we validate an approach to generate genome-wide polymorphism datasets from accessible sample types in the zoonotic hookworm (Ancylostoma ceylanicum) then apply the validated approach in the human hookworm (Necator americanus) to compare laboratory- and field-derived samples. We first present an optimized method for purifying nucleic acid from individual third-stage hookworm larvae (L3s). We then measure the accuracy of variant call datasets generated through whole genome amplification (WGA) and next-generation sequencing (NGS). We demonstrate that WGA via multiple displacement amplification (MDA) introduces predictable biases that are exacerbated by low inputs and poor sample preservation but show that with sufficient input mass ([≥]0.1ng) we are still able to produce highly accurate variant call datasets from nucleic acid concentrations that reflect those of individual L3s. Using our validated approach, we infer laboratory- and field-collected samples of N. americanus as distinct populations, with higher levels of heterozygosity and nucleotide diversity identified in field-collected samples, suggesting signatures of inbreeding and/or drift are detectable in laboratory specimens within several years of initiation of infection. We also show that, despite expected reductions in heterozygosity, laboratory samples still possess numerous heterozygous sites, and we demonstrate that a reference genome generated from an adult worm from an early laboratory passage performs well for variant calling in both laboratory- and field-derived samples. Moving forward, our optimized method for nucleic acid purification can be broadly applied to generate input for any amplification-based approach where sequencing individual hookworm L3s, rather than a pool of specimens, is preferred. Our validated population genomics workflow can be used to characterize structure and connectivity of hookworm populations in endemic communities, with the goal of leveraging these insights to improve our approach to hookworm treatment and control. AUTHOR SUMMARYInfection with the human hookworm, Necator americanus, is a significant cause of morbidity in the global south. Population genomic techniques have the potential to improve our understanding of hookworm infection and control. However, third-stage larvae, or L3, which are the hookworm life-cycle stage we routinely have access to when screening infected people, are less than a millimeter long, and this small size makes generating genomic datasets from individual worms difficult. Here, we introduce an optimized approach for purifying DNA from individual L3s and validate its use with whole genome amplification (WGA) for population genomics. We found that WGA from L3s produces sequencing datasets that are biased in their breadth and depth of coverage across the hookworm genome. But by ensuring a minimum threshold for the mass of DNA input to WGA and setting strict criteria for variant filtration, these biases can be overcome to produce highly accurate variant calls genome-wide. We then use this approach to demonstrate reduced genetic diversity in a recently established laboratory hookworm strain as compared to field-collected samples. Our study outlines a specimen-through-analysis workflow that can be used with accessible sample types to measure population structure and diversity of hookworms in endemic communities.

genomics↗

Genomic Characterization of a Severe West Nile Virus Transmission Season using a Single Reaction Amplicon Sequencing Approach

West Nile virus (WNV) is an endemic arthropod-borne virus that has routinely caused seasonal outbreaks in the United States since it was first detected in 1999. While phylogenetic studies have shown how WNV has diversified and undergone genotype replacement since introduction, more geographically focused studies are needed to understand intricate transmission dynamics at local and regional scales. In this study, we validate the IDT xGen WNV panel, a novel single reaction amplicon-based Next-Generation Sequencing approach, to generate high-quality WNV genomes and compare it to the "Primal Scheme" assay for WNV, a common amplicon sequencing strategy. We show that the IDT xGen WNV panel generated complete and accurate WNV genomes and was more robust to amplicon drop out compared to the current sequencing approaches. Additionally, we used this approach to generate 100 complete WNV genomes from surveillance pools of mosquitoes collected in Nebraska during the 2023 outbreak. Our discrete phylogeographic analysis revealed substantial genetic diversity in WNV genomes from 2023 with minimal clustering across the state. This study demonstrated the utility of a single reaction amplicon-based sequencing approach to generate quality WNV genomes from routine surveillance samples and characterize WNV transmission dynamics in a high-incidence setting.

genomics↗

Generating reference-quality de novo genome assemblies from parasitic nematodes using the Oxford Nanopore Technologies MinION sequencing platform

Parasitic nematode infections represent a significant burden of disease in impoverished populations. Genomic studies of parasitic nematodes have revealed novel drug and vaccine targets and provided unprecedented insights into parasite biology. A key component of these studies is the availability of high-quality reference genomes (i.e., genomes that are contiguous, complete, and accurate) that capture the biological diversity contained within species. However, relatively few genomic resources exist for parasitic nematodes and few species are represented by more than a single reference genome. Streamlined laboratory and computational workflows to generate high-quality reference genomes from individual specimens using a single data source have the potential to increase the availability of genomic data from parasitic nematodes. The Oxford Nanopore Technologies (ONT) MinION is an accessible sequencing platform capable of generating ultra-long read data ideal for assembling genomes. However, lower read-level accuracy of ONT data has previously required assemblies to be error-corrected with more accurate short-read data. In this study, we assessed the quality of de novo genome assemblies for three species of parasitic nematodes (Brugia malayi, Trichuris trichiura, and Ancylostoma caninum) generated using only ONT MinION data. Assemblies were benchmarked against current reference genomes and against additional assemblies that were supplemented with short-read Illumina data through polishing or hybrid assembly approaches. For each species, assemblies generated using only MinION data had similar or superior measures of contiguity, completeness, and gene content. In terms of gene composition, depending on the species, between 88.9-97.6% of complete coding sequences predicted in MinION data only assemblies were identical to those predicted in assemblies polished with Illumina data. Polishing MinION data only assemblies with Illumina data therefore improved gene-level accuracy to a degree. Furthermore, modified DNA extraction and library preparation protocols produced sufficient genomic DNA from B. malayi and T. trichiura to generate de novo assemblies from individual specimens. Data Availability StatementQuality-controlled MinION and Illumina data for each species are deposited in the NCBI Sequence Read Archive (SRA) under the BioProject accession ID PRJNA1074771 under BioSample accession nos. SAMN39888962 (Brugia malayi), SAMN39888963 (Trichuris trichiura) and SAMN39888964 (Ancylostoma caninum). Final assemblies for each species are publicly available on GenBank.

genomics↗