bioRxiv ScienceSearch

Biology subjects

Organick, L. W.

Publications and source records attributed to Organick, L. W..

1 recordsLinked to original sources

Experimental Assessment of PCR Specificity and Copy Number for Reliable Data Retrieval in DNA Storage

Synthetic DNA has been gaining momentum as a potential storage medium for archival data storage1-9. Digital information is translated into sequences of nucleotides and the resulting synthetic DNA strands are then stored for later individual file retrieval via PCR7-9 (Fig. 1a). Using a previously presented encoding scheme9 and new experiments, we demonstrate reliable file recovery when as few as 10 copies per sequence are stored, on average. This results in density of about 17 exabytes/g, nearly two orders of magnitude greater than prior work has shown6. Further, no prior work has experimentally demonstrated access to specific files in a pool more complex than approximately 106 unique DNA sequences9, leaving the issue of accurate file retrieval at high data density and complexity unexamined. Here, we demonstrate successful PCR random access using three files of varying sizes in a complex pool of over 1010 unique sequences, with no evidence that we have begun to approach complexity limits. We further investigate the role of file size on successful data recovery, the effect of increasing sequencing coverage to aid file recovery, and whether DNA strands drop out of solution in a systematic manner. These findings substantiate the robustness of PCR as a random access mechanism in complex settings, and that the number of copies needed for data retrieval does not compromise density significantly.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC=\"FIGDIR/small/565150_fig1.gif\" ALT=\"Figure 1\">\nView larger version (29K):\norg.highwire.dtl.DTLVardef@4b79a7org.highwire.dtl.DTLVardef@11fca45org.highwire.dtl.DTLVardef@18b830org.highwire.dtl.DTLVardef@e47101_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO (a) A high-level representation of the DNA data storage pipeline. (b) (Left) The bar chart depicts contents of the initial, undiluted pool. (Right) The illustration shows the serial nature of subsequent dilutions. Mean copy number refers to the mean number of copies of each files unique sequences as determined by qPCR (Supplemental Section 1). One serial dilution used water as the diluent in each step; the other used a solution of 150Nmers to dilute the pool to much greater complexity. (c) Details of how the samples were diluted. Note that the dilution steps were identical regardless of diluent. The smallest percent of pool accessed is calculated by dividing the size of the smallest file by the number of unique sequences in the 1 {micro}L of solution used for PCR random access. This percentage refers to the 150N diluent pool since the small file in the water diluent pool is a constant 0.13%.\n\nC_FIG

synthetic biology