bioRxiv Science⌕ Search

Biology subjects

Zufan, S. E.

Publications and source records attributed to Zufan, S. E..

3 recordsLinked to original sources

Bridging genomic gaps: A versatile SARS-CoV-2 benchmark dataset for adaptive laboratory workflows

Genomic sequencings adoption in public health laboratories (PHLs) for pathogen surveillance is innovative yet challenging, particularly in the realm of bioinformatics. Low- and middle-income countries (LMICs) face increased difficulties due to supply chain volatility, workforce training, and unreliable infrastructure such as electricity and internet services. These challenges also extend to high-income countries (HICs) where bioinformatics is nascent in PHLs and hampered by a lack of specialized skills and computational infrastructure. This underlines the urgency for flexible and resource-aware strategies in genomic sequencing to improve global pathogen surveillance. In response to these challenges, the present research was conducted to identify and analyse key variables influencing the quality and accuracy of amplicon sequence data. An extensive benchmark dataset was developed that encompassed a diverse collection of isolates, viral loads, primer schemes, library preparation methods, sequencing technologies, and basecalling models, totalling 750 sequences. This dataset was analysed with bioinformatic workflows selected for varying levels of technical capacity. The evaluation focused on quality metrics, consensus accuracy, and common genomic epidemiological indicators. The analysis uncovers complex interactions between multiple parameters in laboratory and bioinformatic processes. emphasising resource-constrained PHLs, practical guidelines are proposed. Insights from the benchmark dataset aim to guide the establishment of specific laboratory and bioinformatics protocols for amplicon sequencing in these settings. The findings can also be used to guide the creation of specialised training curricula, further advancing genomic equity. The benchmark dataset itself allows laboratories to customise and evaluate workflows, catering to their distinct requirements and capacities. Such a holistic approach is imperative to build the capacity to monitor pathogens worldwide. Author summaryThis study marks a step toward equity in the field of pathogen genomics, especially for resource-constrained PHLs. It develops and evaluates a comprehensive amplicon sequencing benchmark dataset, offering vital insights for PHLs engaged in genomic surveillance. In particular, the study finds that the choice of basecaller model has a minimal impact on the quality and accuracy of consensus sequences derived from ONT data, which is crucial for labs with limited computational resources. It also highlights the effectiveness of longer amplicons in ensuring consistent coverage and reducing amplicon dropouts at higher viral loads. While Illumina remains a gold standard for data quality, the combination of the Midnight primer scheme with ONTs Rapid library preparation is shown to be a viable alternative, reducing costs, procedural complexity, and hands-on time. The study synthesises these findings into practical guidelines to aid in the development of amplicon sequencing workflows for SARS-CoV-2 with implications for other pathogens.

genomics↗

High performance enrichment-based genome sequencing to support the investigation of hepatitis A virus outbreaks

Hepatitis A virus (HAV) infections are an increasing public health concern in low-endemicity regions due to outbreaks from foodborne infections and sustained transmission among vulnerable groups, including persons experiencing homelessness, those who inject drugs, and men who have sex with men (MSM), which is further compounded by aging, unvaccinated populations. DNA sequence characterisation of HAV for source tracking is performed by comparing small subgenomic regions of the virus. While this approach has been successful when robust epidemiological data are available, poor genetic resolution can lead to conflation of outbreaks with sporadic cases. HAV outbreak investigations would greatly benefit from the additional phylogenetic resolution obtained by whole virus genome sequence comparisons. However, HAV genomic approaches can be difficult because of challenges in isolating the virus, low sensitivity of direct metagenomic sequencing in complex sample matrices like various foods such as fruits, vegetables and molluscs, and difficulty designing highly multiplexed PCR primers across diverse HAV genotypes. Here, we introduce a proof-of-concept pan-HAV oligonucleotide hybrid capture enrichment assay from serum and frozen berry specimens that yields complete and near-complete HAV genomes from as few as four input HAV genome copies. We used this method to recover HAV genomes from human serum specimens with high C{tau} values (34{middle dot}7--42{middle dot}7), with high assay performance for all six human HAV sub-genotypes, both contemporary and historical. Our approach provides a highly sensitive and streamlined workflow for HAV WGS from diverse sample types, that can be the basis for harmonised and high-resolution molecular epidemiology during HAV outbreak surveillance. ImportanceThis proof-of-concept study introduces a hybrid capture oligo panel for whole genome sequencing (WGS) of all six human pathogenic hepatitis A virus (HAV) subgenotypes, exhibiting a higher sensitivity than some conventional genotyping assays. The ability of hybrid capture to enrich multiple targets allows for a single, streamlined workflow, thus facilitating the potential harmonization of molecular surveillance of HAV with other enteric viruses. Even challenging sample matrices can be accommodated, making it suitable for broad implementation in clinical and public health laboratories. The ability to capture small amounts of virus from complex samples is promising for passive surveillance application to environmental substrates, such as wastewater. This innovative approach has significant implications for enhancing multijurisdictional outbreak investigations, as well as our understanding of the global diversity and transmission dynamics of HAV.

genomics↗

Bioinformatic investigation of discordant sequence data for SARS-CoV-2: insights for robust genomic analysis during pandemic surveillance

The COVID-19 pandemic has necessitated the rapid development and implementation of whole genome sequencing (WGS) and bioinformatic methods for managing the pandemic. However, variability in methods and capabilities between laboratories has posed challenges in ensuring data accuracy. A national working group comprising 18 laboratory scientists and bioinformaticians from Australia and New Zealand was formed to improve data concordance across public health laboratories (PHLs). One effort, presented in this study, sought to understand the impact of methodology on consensus genome concordance and interpretation. Data were retrospectively obtained from the 2021 Royal College of Pathologists of Australasia Quality Assurance Programs (RCPAQAP) SARS-CoV-2 WGS proficiency testing program (PTP), which included 11 participating Australian laboratories. The submitted consensus genomes and reads from eight contrived specimen were investigated, focusing on discordant sequence data, and findings were presented to the working group to inform best practices. Despite using a variety of laboratory and bioinformatic methods for SARS-CoV-2 WGS, participants largely produced concordant genomes. Two participants returned five discordant sites in a high Ct replicate which could be resolved with reasonable bioinformatic quality thresholds. We noted ten discrepancies in genome assessment that arose from nucleotide heterogeneity at three different sites in three cell-culture derived control specimen. While these sites were ultimately accurate after considering the participants bioinformatic parameters, it presented an interesting challenge for developing standards to account for intrahost single nucleotide variation (iSNV). Observed differences had little to no impact on key surveillance metrics, lineage assignment and phylogenetic clustering, while genome coverage <90% affected both. We recommend PHLs bioinformatically generate two consensus genomes with and without ambiguity thresholds for quality control and downstream analysis, respectively, and adhere to a minimum 90% genome coverage threshold for inclusion in surveillance interpretations. We also suggest additional PTP assessment criteria, including primer efficiency, detection of iSNVs, and minimum genome coverage of 90%. This study underscores the importance of multidisciplinary national working groups in informing guidelines in real time for bioinformatic quality acceptance criteria. It demonstrates the potential for enhancing public health responses through improved data concordance and quality control in SARS-CoV-2 genomic analysis during pandemic surveillance. Data summaryThe authors confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. Impact statementAmidst the COVID-19 pandemic, a unique collaboration between a national multidisciplinary working group and a quality assurance program facilitated ongoing development of standardized quality control criteria and analysis methods for high-quality SARS-CoV-2 genomic approaches across Australia. With this article, we shed light on the robustness of amplicon sequencing and analysis methods to produce highly concordant genomes, while also presenting additional assessment criteria to guide laboratories in identifying areas for improvement. Insights from this nationwide collaboration underscore the need for real-time knowledge-sharing and iterative refinements to quality standards, particularly as situations and methods evolve during a pandemic. While the spotlight is on SARS-CoV-2, the analyses and findings have universal implications for genomic surveillance during infectious disease outbreaks. As WGS becomes increasingly central in outbreak surveillance, continuous evaluation and collaboration, like that described here, are vital to ensure data accuracy and inform future public health responses.

genomics↗