bioRxiv ScienceSearch

Biology subjects

Ellington, A.

Publications and source records attributed to Ellington, A..

2 recordsLinked to original sources

Evolving Bacterial Fitness with an Expanded Genetic Code

Evolution has for the most part used the canonical 20 amino acids of the natural genetic code to construct proteins. While several theories regarding the evolution of the genetic code have been proposed, experimental exploration of these theories has largely been restricted to phylogenetic and computational modeling. The development of orthogonal translation systems has allowed noncanonical amino acids to be inserted at will into proteins. We have taken advantage of these advances to evolve bacteria to accommodate a 21 amino acid genetic code in which the amber codon ambiguously encodes either 3-nitro-L-tyrosine or stop. Such an ambiguous encoding strategy recapitulates numerous models for genetic code expansion, and we find that evolved lineages first accommodate the unnatural amino acid, and then begin to evolve on a neutral landscape where stop codons begin to appear within genes. The resultant lines represent transitional intermediates on the way to the fixation of a functional 21 amino acid code.

evolutionary biology

A highly parallel strategy for storage of digital information in living cells

Encoding arbitrary digital information in DNA has attracted attention as a potential avenue for large scale and long term data storage. However, in order to enable DNA data storage technologies there needs to be improvements in data storage fidelity (tolerance to mutation), the facility of writing and reading the data (biases and systematic error arising from synthesis and sequencing), and overall scalability. To this end, we have developed and implemented an encoding scheme that is suitable for detecting and correcting errors that may arise during storage, writing, and reading, such as those arising from nucleotide substitutions, insertions, and deletions. We propose a scheme for parallelized long term storage of encoded sequences that relies on overlaps rather than the address blocks found in previously published work. Using computer simulations, we illustrate the encoding, sequencing, decoding, and recovery of encoded information, ultimately demonstrating the possibility of a successful round-trip read/write. These demonstrations show that in theory a precise control over error tolerance is possible. Even after simulated degradation of DNA, recovery of original data is possible owing to the error correction capabilities built into the encoding strategy. A secondary advantage of our method is that the statistical characteristics (such as repetitiveness and GC-composition) of encoded sequences can also be tailored without sacrificing the overall ability to store large amounts of data. Finally, the combination of the overlap-based partitioning of data with the LZMA compression that is integral to encoding means that the entire sequence must be present for successful decoding. This feature enables inordinately strong encryptions. As a potential application, an encrypted pathogen genome could be could be distributed and carried by cells without danger of being expressed, and could not even be read out in the absence of the entire DNA consortium.

synthetic biology