Skip to main content
GES Center Lectures, NC State University
Genetic Engineering & Society Center | Where biotechnology meets society—ethics, policy, and practice.

S14E5 - James Tuck - DNA-based data storage: How to store a library in a raindrop

Oct. 6, 2026 GES Colloquium | An introduction to the basics of DNA-based data storage and some recent advances from Dr. Tuck’s group.

How to Store a Library in a Raindrop: DNA-based Data Storage

James Tuck, PhD | Website

Professor and Interim Department Head in Electrical and Computer Engineering, NC State University  

Full details at https://ges.research.ncsu.edu/event/colloquium-2026-10-06/ | Watch the video

Related links:

  • Kyle J. Tomek, Kevin Volkel, Alexander Simpson, Austin G. Has, Elaine W. Indermaur, James M. Tuck, Albert J. Keung. Driving the Scalability of DNA-Based Information Storage Systems, ACS Synth. Biol.2019861241-1248, https://doi.org/10.1021 

Zoom Summary

Overview

This colloquium presentation, delivered by James Tuck, Professor and Interim Department Head in the Department of Electrical and Computer Engineering at NC State University, explored the emerging field of DNA-based data storage. The talk addressed the fundamental question of whether an entire library's worth of data could be stored in a single raindrop, and provided a comprehensive overview of how DNA can serve as an ultra-dense, long-lasting storage medium for digital information.

Key Concepts or Theories

  • Binary-to-DNA Mapping: Digital data (zeros and ones) can be systematically converted into DNA base sequences (A, C, G, T), with each base representing two bits of information.
  • Data Temperature Hierarchy: Data is classified as hot, warm, cold, or frozen based on access frequency, with DNA storage best suited for cold/frozen archival data.
  • Six-Step DNA Storage Pipeline: The process of storing and retrieving data in DNA involves encoding, synthesis, storage, access, sequencing, and decoding.
  • GC Balance and Homopolymer Avoidance: Effective DNA storage requires careful sequence design to maintain chemical stability and sequencing accuracy.
  • Error Correction: Redundancy and mathematical coding techniques are used to ensure reliable data recovery despite synthesis and sequencing errors.
  • Molecular Computing: Enzymatic and chemical reaction networks can potentially enable computation directly on stored DNA, reducing the need to transfer data back to digital systems.

Important Questions Raised

  • Can DNA storage be made cost-effective enough for widespread commercial adoption?
  • How do existing data privacy regulations (e.g., the EU's right to erasure) apply when personal data is stored in a DNA archive that cannot be selectively unmixed?
  • What knowledge should humanity choose to preserve in long-lasting DNA archives, and who controls that decision?
  • Can living biological systems, such as yeast, be harnessed to store and maintain digital data?
  • How do DNA strand length, synthesis cost, and sequencing error rates interact to define optimal storage parameters?

Key Takeaways and Summary of Learning Objectives

  • DNA offers extraordinary storage density, with a theoretical peak of approximately 450 exabytes per gram, far exceeding current digital storage technologies.
  • The global data sphere is estimated at around 200 zettabytes and is growing exponentially, creating urgent demand for new storage solutions.
  • DNA storage is most practical for cold or frozen data — information that must be retained for decades or centuries but accessed infrequently.
  • A six-step pipeline (encode, synthesize, store, access, sequence, decode) forms the foundation of any DNA-based storage system.
  • Error correction strategies, borrowed from decades of computer engineering research, can ensure reliable data recovery even in the presence of synthesis and sequencing errors.
  • Commercialization of DNA storage is underway, with companies such as BioMemory and Atlas Biosciences actively developing the technology.
  • Significant challenges remain, including reducing synthesis costs, scaling archive sizes beyond the current ~200 megabyte laboratory record, and integrating DNA storage with existing IT infrastructure.
  • DNA storage raises important ethical, legal, and policy questions around data ownership, deletion rights, and the long-term stewardship of human knowledge.

Topic 1: The Case for DNA-Based Data Storage

The exponential growth of global data — estimated at approximately 200 zettabytes today and projected to continue rising — is straining the capacity, cost, and energy efficiency of conventional storage technologies. Hard drives, magnetic tape, CDs, and flash memory all have finite lifespans ranging from a few years to a few decades, and none can match the density that DNA theoretically offers.

James Tuck introduced DNA as a compelling alternative by drawing on a straightforward back-of-the-envelope calculation: a human body, with roughly 30 trillion cells each containing approximately 3 billion base pairs, could theoretically hold around 22.5 zettabytes of information — comparable to the entire global data sphere. At the molecular level, the peak theoretical storage density of DNA is approximately 450 exabytes per gram. Translating this to a practical example, a single raindrop (approximately 50 microliters) could potentially hold around one exabyte of data, equivalent to roughly 56 copies of the Hunt Library at NC State.

Beyond density, DNA offers exceptional longevity. Researchers have successfully sequenced DNA from fossils millions of years old, and laboratory aging studies suggest that properly preserved synthetic DNA could remain readable for centuries or longer — far surpassing the 5–30 year lifespans of current storage media. Additionally, DNA is likely to remain permanently relevant to humanity as long as we are biological beings, unlike obsolete formats such as floppy disks or cassette tapes.

Within the data storage hierarchy, DNA is best positioned as a medium for cold or frozen data: information that must never be deleted but is accessed very rarely, such as medical records, government archives, research datasets, financial records, and the broader corpus of human knowledge.

Relevant Q\&A

Question: What are the density advantages of DNA compared to conventional storage, and how does this translate to a practical example like a library?
Answer: At a theoretical peak density of approximately 450 exabytes per gram, a single raindrop of DNA-dissolved water could store around one exabyte of data — enough to hold approximately 56 copies of the Hunt Library. This density advantage stems from the fact that each DNA base can encode two bits of information, and DNA molecules can be packed extremely tightly in solution.


Topic 2: How DNA Data Storage Works — The Six-Step Pipeline

James Tuck outlined a six-step process that forms the operational backbone of any DNA-based storage system.

Step 1 — Encode: Any digital file, regardless of format, is broken into small chunks (typically 20–40 bytes each) that can fit on a single short DNA strand. Each chunk is assigned an index to enable later reassembly, and a file identifier is added so that multiple files can coexist in the same archive. The binary data is then converted into DNA base sequences. Not all sequences are suitable: GC balance (roughly equal proportions of G/C and A/T bases) must be maintained, and homopolymer runs (e.g., long stretches of the same base) must be avoided to ensure sequencing accuracy. As a result, practical systems achieve approximately 1 to 1.5 bits per base rather than the theoretical maximum of 2 bits per base. Error correction codes are also embedded to enable recovery from synthesis and sequencing errors.

Step 2 — Synthesize: The designed sequences are manufactured into physical DNA, either by sending sequence files to commercial providers such as IDT or Twist Biosciences, or using in-house DNA synthesis equipment. Current bulk oligosynthesis costs are approximately one-tenth of a cent per base, making large-scale synthesis expensive but improving. Strand lengths of 100–300 bases represent a practical and cost-effective sweet spot.

Step 3 — Store: The synthesized DNA strands are stored in test tubes or arrays of tubes. A single tube can potentially hold between a terabyte and a petabyte of data. Multiple tubes can be organized into racks to form a large-scale archive.

Step 4 — Access: Because molecules cannot be unmixed once combined, selective file retrieval relies on Polymerase Chain Reaction (PCR). Unique primer sequences corresponding to each file's identifier are used to amplify only the target file's molecules, making them overwhelmingly abundant relative to the rest of the archive before sequencing.

Step 5 — Sequence: The amplified DNA is read by a sequencing machine, producing base-called sequences and quality scores. Both high-throughput sequencing and nanopore sequencing can be used, though nanopore sequencing introduces higher error rates that must be accounted for in the encoding design.

Step 6 — Decode: Software processes the raw sequencing reads to filter by file ID, cluster repeated reads, correct base errors, fill in missing data, and reassemble the chunks in their original order to reconstruct the original digital file. All decoding is performed in silico.

Relevant Q\&A

Question: Do the error correction sections account for mutations and errors that accumulate over long storage periods, and can data integrity be guaranteed to shareholders?
Answer: Yes. Error correction is a well-established field dating back nearly a century. As long as the system's worst-case error rate is characterized, redundancy and mathematical coding techniques can be designed to reliably recover the original data. PCR naturally produces many copies, which aids error correction, and more sophisticated mathematical codes provide even stronger guarantees. Provided the system operates within its nominal parameters, a strong case can be made to shareholders that data will always be recoverable.

Question: Is the decoding process performed in silico or in vitro?
Answer: All decoding and translation steps are performed in silico. The sequencer produces base-called reads, and software handles all subsequent processing to reconstruct the original file.

Question: What is the optimal strand length for DNA storage?
Answer: Current studies have focused on strands of approximately 100 to 300 bases, which represent a balance between synthesis cost, chemical stability, and PCR amplification efficiency. More detailed quantitative models of how reliability changes with strand length are still needed before a definitive optimal length can be specified.


Topic 3: Advanced Directions — Molecular Computing and Hybrid Systems

James Tuck discussed a significant bottleneck in DNA storage systems: the memory bottleneck. Even if vast amounts of data are stored in DNA, retrieving and processing it requires sequencing, decoding, and loading the data into a conventional computer — a slow and computationally expensive process. If the desired data is not found, the entire cycle must be repeated.

To address this, James Tuck proposed the concept of molecular computing: writing programs using enzymes or chemical reaction networks that can operate directly on the DNA pool in parallel, returning a single readable output without requiring full retrieval and digital processing. This approach could dramatically reduce latency and energy consumption for certain query types.

Looking further ahead, James Tuck described a vision of hybrid computing systems that integrate a conventional digital side (processors, algorithms, AI) with a molecular side capable of reading, writing, and computing on DNA and RNA. Such systems could enable faster experimental cycles, novel forms of molecular sensing, and more seamless coupling between biological and digital information processing.

James Tuck also described a collaborative study with Dr. Keung in NC State's Department of Chemical and Biomolecular Engineering, in which data was stored in living yeast cells. The yeast were engineered to display surface proteins indicating which file was stored inside, enabling cell-sorting technology to locate and retrieve specific files — demonstrating that biological systems can be harnessed to add functional features to DNA storage architectures.

Relevant Q\&A

Question: Can synthetic or artificial DNA bases beyond A, T, G, and C be used to increase storage density?

Answer: Yes. The standard four-base mapping represents a lower bound on what is achievable. Synthetic bases, chemical modifications such as methylation, and other non-standard nucleotides can all encode additional information per position. The key requirement is that compatible synthesis and sequencing technologies exist to write and read those bases reliably. Nanopore sequencing, for example, could potentially be adapted to detect synthetic bases through their distinct electrical signals.

Question: Could data be stored in living organisms, and what are the implications?

Answer: It is technically feasible. A study conducted with Dr. Keung at NC State demonstrated data storage in yeast, with cells engineered to display proteins indicating their stored file, enabling retrieval via cell sorting. While living systems introduce biological complexity and variability, they also offer self-replication and maintenance capabilities that could be advantageous. The field of computing on information in living cells is active and growing, though James Tuck's own focus remains on purely synthetic systems for reliability and consistency.


Topic 4: Applications, Commercialization, and Ethical Considerations

James Tuck outlined the most practical near-term applications for DNA storage, all characterized by the need for long-term retention and infrequent access:

  • Medical records: Legal requirements mandate long retention periods; DNA's density means archives need not rely on cloud infrastructure.
  • Government archives: Governments accumulate vast, sensitive records over timescales exceeding individual human lifespans.
  • Research data: Telescopes, genomic studies, and large instruments continuously generate data that researchers wish to retain indefinitely.
  • Financial records and legal documents: Long-term retention requirements align well with DNA's durability.
  • Archiving human knowledge: Books, languages, art, science, and history could be encoded in a form compact and durable enough to survive for millennia, and potentially transported beyond Earth.

Two companies actively commercializing DNA storage were highlighted: BioMemory (Europe) and Atlas Biosciences (a Twist Biosciences spin-out), both working to reduce synthesis costs and build out the full DNA storage ecosystem.

James Tuck also raised important ethical and policy challenges. Current EU regulations and California's Consumer Privacy Act grant citizens the right to have their personal data deleted. However, once personal data is mixed into a DNA archive, it cannot be selectively removed — raising unresolved questions about compliance, archive destruction, and data governance. Additionally, scaling DNA storage to commercial viability would require producing approximately 9 trillion DNA bases per day (equivalent to roughly 3,000 human genome equivalents), far exceeding current global synthesis capacity, necessitating major advances in parallel synthesis technologies.

Relevant Q\&A

Question: How would data deletion rights (e.g., EU right to erasure) be handled in a DNA archive where molecules cannot be unmixed?

Answer: This remains an open and unresolved policy question. Once personal data is incorporated into a mixed DNA archive, selective removal is effectively impossible. Potential approaches include never storing personally identifiable information in DNA archives, or destroying the entire archive — neither of which is ideal. James Tuck acknowledged this as a significant challenge that the field has not yet solved.


Actionable Next Steps

Next week's GES Colloquium will be held online via Zoom. No in-person attendance is required, though participants may join from the usual room if preferred. The speaker will be Ed Perry, a professor in the Economics Department at Iowa State University, presenting research on the rural health effects of pesticide exposure and genetically engineered crop adoption on U.S. farmland.

__ Recorded from NC State’s GES Colloquium, this podcast examines how biotechnologies take shape in the world: microbiome engineering in built environments, gene editing and gene drives, forest and agricultural genomics, data governance and equity, risk and regulation, sci-art, and public engagement in practice.

Genetic Engineering and Society Center

Colloquium Home | Zoom Registration | Watch Colloquium Videos | LinkedIn | Newsletter

GES Center at NC State University—Integrating scientific knowledge & diverse public values in shaping the futures of biotechnology.

Produced by Patti Mulligan, Communications Director, GES Center, NC State

Find out more at https://ges-center-lectures-ncsu.pinecast.co