bioRxiv ScienceSearch

bioRxiv · 10.1101/565150

Experimental Assessment of PCR Specificity and Copy Number for Reliable Data Retrieval in DNA Storage

Abstract

Synthetic DNA has been gaining momentum as a potential storage medium for archival data storage1-9. Digital information is translated into sequences of nucleotides and the resulting synthetic DNA strands are then stored for later individual file retrieval via PCR7-9 (Fig. 1a). Using a previously presented encoding scheme9 and new experiments, we demonstrate reliable file recovery when as few as 10 copies per sequence are stored, on average. This results in density of about 17 exabytes/g, nearly two orders of magnitude greater than prior work has shown6. Further, no prior work has experimentally demonstrated access to specific files in a pool more complex than approximately 106 unique DNA sequences9, leaving the issue of accurate file retrieval at high data density and complexity unexamined. Here, we demonstrate successful PCR random access using three files of varying sizes in a complex pool of over 1010 unique sequences, with no evidence that we have begun to approach complexity limits. We further investigate the role of file size on successful data recovery, the effect of increasing sequencing coverage to aid file recovery, and whether DNA strands drop out of solution in a systematic manner. These findings substantiate the robustness of PCR as a random access mechanism in complex settings, and that the number of copies needed for data retrieval does not compromise density significantly.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC=\"FIGDIR/small/565150_fig1.gif\" ALT=\"Figure 1\">\nView larger version (29K):\norg.highwire.dtl.DTLVardef@4b79a7org.highwire.dtl.DTLVardef@11fca45org.highwire.dtl.DTLVardef@18b830org.highwire.dtl.DTLVardef@e47101_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO (a) A high-level representation of the DNA data storage pipeline. (b) (Left) The bar chart depicts contents of the initial, undiluted pool. (Right) The illustration shows the serial nature of subsequent dilutions. Mean copy number refers to the mean number of copies of each files unique sequences as determined by qPCR (Supplemental Section 1). One serial dilution used water as the diluent in each step; the other used a solution of 150Nmers to dilute the pool to much greater complexity. (c) Details of how the samples were diluted. Note that the dilution steps were identical regardless of diluent. The smallest percent of pool accessed is calculated by dividing the size of the smallest file by the number of unique sequences in the 1 {micro}L of solution used for PCR random access. This percentage refers to the 150N diluent pool since the small file in the water diluent pool is a constant 0.13%.\n\nC_FIG

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Organick, L. W., Chen, Y.-J., Dumas Ang, S., Lopez, R., Strauss, K., Ceze, L.. 2019-03-04. Experimental Assessment of PCR Specificity and Copy Number for Reliable Data Retrieval in DNA Storage. https://doi.org/10.1101/565150

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Living electronic transistors with tunable conductivity

Electroactive bacteria, like Shewanella oneidensis, can couple the oxidation of organic electron donors to the reduction of external conductive surfaces, such as minerals and electrodes, by utilizing multiheme cytochromes to carry charge from within the cell to external surfaces. Additionally, multiheme cytochromes facilitate gateable, long-distance (micrometer-scale) redox conduction along the outer membrane and across multiple cells bridging electrodes. While electroactive microbes are being used to develop bioelectrochemical devices, there have been limited efforts to use synthetic biology to exert additional control over microbes serving as device components. Thus, this work implements an optogenetic biofilm patterning gene circuit and a small molecule sensor in S. oneidensis to simultaneously control cell deposition and cytochrome expression. This allows for photolithographic patterning of biofilms possessing tunable electrical properties controlled with small molecules. This system demonstrates tunable electrochemical activity, redox conduction, intrinsic biofilm conductivity, and negative differential transconductance as a function of cytochrome expression. Additionally, temperature-dependent measurements of this tunable biofilm conduction reveal changes in activation energy as a function of cytochrome expression. Through this combination of synthetic biology and electrochemistry, simultaneous control over biofilm geometry and conductivity sheds light on fundamental microbial electron transport processes, and it enables the construction of living electronic devices.

synthetic biology

Evolutionary stabilisation of stressful metabolism via integrated biocomputing and essential-gene metabolic locking circuits

Synthetic genetic circuits enable microbial differentiation from growth to production, yet metabolic burden, imbalance and toxicity frequently drive strain degeneration. Yeast strains engineered to produce different terpene products exhibited divergent genetic responses to metabolic stresses, but commonly underwent progressive loss of induction of synthetic GAL regulatory circuits, either across the entire population or within subpopulations. Using di- and tri-input biocomputing circuits, the essential glutamine synthetase gene GLN1 was coupled to GAL induction, thereby enabling stabilisation and evolutionary adaptation of the synthetic genetic circuits and stressful heterologous terpene synthetic pathways. The integrated biocomputing and metabolic coupling circuit systems not only prevent strain degeneration but also enable interrogation of non-degenerative evolutionary shifts, providing a platform for metabolic engineering optimisation.

synthetic biology

Unbiased and scalable reduction of diverse bacterial genomes

The genome is a complex, integrated system where the functions and regulatory interactions of its many components remain poorly understood. Genome minimization aims to reduce genomic complexity by removing non-essential elements to reveal the fundamental building blocks of cellular life. However, current minimization strategies are often slow and species-specific due to a reliance on prior information, and limited to producing single, isolated strains, which obscures the diverse ways a genome can adapt to large-scale DNA removal. Here we show the development and application of Stochastic Lineage-based Iterative Minimization (SLIM) a modular, high-throughput platform for unbiased genome reduction across phylogenetically diverse bacteria. We apply SLIM to generate a library of genome-reduced Escherichia coli lineages. We then interrogate the lineages, identifying both universal and lineage-specific transcriptional and translational reprogramming in response to deletions. We demonstrate that these expression dynamics drive environment-dependent fitness, allowing us to pinpoint a single gene deletion in one genome-reduced lineage as the driver of a measurable environmental growth defect. Beyond E. coli, we successfully deploy SLIM in phylogenetically distinct bacterial taxa to rapidly reduce the genomes of Shigella flexneri and Pseudomonas putida, distinct genus and order respectively from E. coli, without species-specific optimization. Our results establish a scalable, generalizable framework for navigating the vast landscape of minimized genomes, providing a powerful new tool for functional discovery and the rational design of synthetic genomic chassis.

synthetic biology