bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.02.23.529831

Optimized SMRT-UMI protocol produces highly accurate sequence datasets from diverse populations - application to HIV-1 quasispecies

Abstract

Pathogen diversity resulting in quasispecies can enable persistence and adaptation to host defenses and therapies. However, accurate quasispecies characterization can be impeded by errors introduced during sample handling and sequencing which can require extensive optimizations to overcome. We present complete laboratory and bioinformatics workflows to overcome many of these hurdles. The Pacific Biosciences single molecule real-time platform was used to sequence PCR amplicons derived from cDNA templates tagged with universal molecular identifiers (SMRT-UMI). Optimized laboratory protocols were developed through extensive testing of different sample preparation conditions to minimize between-template recombination during PCR and the use of UMI allowed accurate template quantitation as well as removal of point mutations introduced during PCR and sequencing to produce a highly accurate consensus sequence from each template. Handling of the large datasets produced from SMRT-UMI sequencing was facilitated by a novel bioinformatic pipeline, Probabilistic Offspring Resolver for Primer IDs (PORPIDpipeline), that automatically filters and parses reads by sample, identifies and discards reads with UMIs likely created from PCR and sequencing errors, generates consensus sequences, checks for contamination within the dataset, and removes any sequence with evidence of PCR recombination or early cycle PCR errors, resulting in highly accurate sequence datasets. The optimized SMRT-UMI sequencing method presented here represents a highly adaptable and established starting point for accurate sequencing of diverse pathogens. These methods are illustrated through characterization of human immunodeficiency virus (HIV) quasispecies. Author SummaryThere is a great need to understand the genetic diversity of pathogens in an accurate and timely manner, but many errors can be introduced during the sample handling and sequencing steps which may prevent accurate analyses. In some cases, the errors introduced during these steps can be indistinguishable from real genetic variation and prevent analyses from identifying true sequence variation present in the pathogen population. There are established methods which can help to prevent these types of errors, but can involve many different steps and variables, all of which must be optimized and tested together to ensure the desired effect. Here we show results from testing different methods on a set of HIV+ blood plasma samples and arrive at a streamlined laboratory protocol and bioinformatic pipeline which prevents or corrects for different types of errors that can arise in sequence datasets. These methods should be an accessible starting point for anyone wanting accurate sequencing without extensive optimizations.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Westfall, D. H., Deng, W., Pankow, A., Murrell, H., Chen, L., Zhao, H., Williamson, C., Rolland, M. M., Murrell, B., Mullins, J. I.. 2023-02-24. Optimized SMRT-UMI protocol produces highly accurate sequence datasets from diverse populations - application to HIV-1 quasispecies. https://doi.org/10.1101/2023.02.23.529831

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Trans-branching of polyubiquitin chains orchestrates the DNA replication stress response

Polyubiquitin chain geometry dictates functional consequences of ubiquitylation. Although branched polyubiquitin chains are abundant in cells, little is known about their functions. Here we show that branching on the DNA replication factor PCNA, mediated by the ubiquitin-conjugating enzyme UBE2K and involving lysines 63 and 48 of ubiquitin, orchestrates the sequence of events in response to replication stress. By inducing VCP-dependent extraction of PCNA from chromatin, branching promotes re-priming of stalled forks and necessitates a BRCA1-dependent pathway of daughter-strand gap repair. Our study identifies hyper-accumulation of daughter-strand gaps as the mechanistic basis underlying the toxicity of inhibitors of the PCNA-specific isopeptidase, USP1, in BRCA1-deficient cells. Moreover, an unexpected preference of UBE2K to operate in trans suggests a general timing mechanism to organize hierarchies amongst ubiquitin signals.

molecular biology↗

Impaired proteostasis is an early feature of the diabetic heart in humans and mice

Diabetes and obesity increase cardiac lipid levels leading to cardiomyopathy and heart failure. We hypothesized that intermittent fasting would reduce cardiac lipid levels. Surprisingly, intermittent fasting increased myocardial triglyceride content, but rescued mortality and attenuated cardiomyopathy in mice overexpressing cardiomyocyte acyl-CoA synthetase 1 (MHC-ACSL1). Lipid overload caused cardiomyocyte accumulation of polyubiquitinated protein aggregates containing desmin, a scaffolding intermediate filament protein, which intermittent fasting prevented. Furthermore, intermittent fasting reversed elevated myocardial C16:0 ceramide content, and knockdown of ceramide synthase CerS5 and CerS6 reduced palmitate-induced protein aggregation, highlighting a role for C16:0 ceramides in this pathology. Conversely, impairing aggrephagy with cardiomyocyte-specific p62 ablation induced heart failure in mice fed a high-fat diet, with paradoxically reduced cardiac lipid content. Crucially, non-failing diabetic human hearts also exhibited protein aggregate pathology. Taken together, these results demonstrate that impaired proteostasis characterizes cardiomyopathy from cardiac lipid overload and identify a promising new therapeutic target for this condition.

molecular biology↗

Spatial profiling and neurovascular communication in the developing and adolescent cortex following prenatal alcohol exposure

Fetal alcohol spectrum disorders (FASD) constitute a wide range of developmental, cognitive, and behavioral impairments caused by prenatal alcohol exposure (PAE). Although neuronal and vascular consequences of PAE have been studied, how alcohol affects the cerebrovasculature within the framework of the neurovascular unit (NVU) across development remains poorly understood. At minimum, the NVU comprises neurons, astrocyte endfeet, and endothelial cells (ECs), which coordinate to maintain brain homeostasis. Here, we used the NanoString Digital Spatial Profiling platform to characterize spatial transcriptomic data from neurons, astrocytes, and ECs from PAE and saccharin (SAC) control cortices at embryonic day 18 (E18) and postnatal day 28 (P28). Differentially expressed genes were then used for Ingenuity Pathway Analysis (IPA) to identify altered biological pathways and perform comparison analyses across developmental time points, while CellChat was used to infer cell cell communication networks. We uncovered thousands of differentially expressed genes and numerous altered pathways and biological processes in PAE cortices across development. Both IPA and CellChat analyses implicated dysregulation of vascular and extracellular matrix (ECM) remodeling, cell adhesion, and neuroinflammatory signaling. CellChat further predicted the loss of several key bidirectional relationships and altered ligand-receptor interactions among neurovascular cell types at E18 and P28. Overall, these findings identify PAE associated alterations in neurovascular gene expression and intercellular signaling across development, providing potential mechanisms by which PAE may disrupt neurodevelopment.

molecular biology↗