Microsoft, UW demonstrate first fully automated DNA data storage
news.microsoft.com
news.microsoft.com
“Our system’s write-to-read latency is approximately 21 h. The majority of this time is taken by synthesis, viz., approximately 305 s per base, or 8.4 h to synthesize a 99-mer payload and 12 h to cleave and deprotect the oligonucleotides at room temperature. After synthesis, preparation takes an additional 30 min, and nanopore reading and online decoding take 6 min.”
Also the amount of data that can be written is very small in general for oligosynthesis. The most high throughout methods are microarrays by Agilent and others (not used in this paper). You can buy about 32 megabits for $6000 (http://www.customarrayinc.com/oligos_main.htm). The actual cost of synthesis is maybe $2000 as a guess.
So currently DNA storage would be good for small datasets that you would like to store for a long time. Physical density is useless most of the time if you can’t efficiently generate a lot of data to begin with.
We would need several orders of magnitude increase in the write capacity, which is slowly being worked on. Typically people would like to compare synthesis costs with Moore’s law, using Carlson curves ( http://www.synthesis.cc/synthesis/2016/03/on_dna_and_transis... ).
However there hasn’t been as much progress in synthesis as in sequencing. Why? My theory is that Theres not a very big market for DNA synthesis, so the big investments needed haven’t really been there. Maybe storage on dna could be that market, but it would need to show quick and easy wins (eg stepping stones of practicality like what early integrated circuits had).
If you can build small runs (<50base pairs) you can make small mutations to dna, or read particular sections of dna. If you can make dna larger than the average protein (~2000bp) you can invent new proteins from scratch rather than modify existing protein sequences. If you can make dna longer than a plasmid (~10,000bp) you can creatively invent a minimal viable replicable and deliverable unit (a plasmid (bacterial virus)). If you can do millions if base pairs you get to chromosomes and can invent Eucaryotic-transmissable storage.
But until you leap those plateaus, you can likely saturate the intervening market. So even if there’s massive pent up theoretical demand for the wholesale invention of genes, there’s no way to really demonstrate it in the current market.
This inability to estimate demand may make it a tricky spot to invest in.
So the value of each individual variant is relatively low, since you don't know if it will work or not. Plus a lot of other investments have to be made into the toolchain (developing assays for the thing you actually care about is not easy).
In your plateaus, the possibility space explodes with each step function. We are already at the first rung of the ladder of building many 50bp variants. Generally, you'd probably want to be able to conduct an experimental cycle with 1000s of variants. At the upper end, you'd want to pay $5000-10000 for each cycle. There is demand for making protein-length sequences, but the prices the market is willing to pay is probably closer to $10/gene (fully customizable 2kbp) rather than the current price.
Another thing is that not all sequences are created equally. Some sequences will just not synthesize well (maybe there's secondary structure), so to limit costs one ought to limit the number of retries per sequence.
Copying isn’t as interesting as de novo synthesis, being able to specify an arbitrary sequence and have it synthesized.
""DNA can store digital information in a space that is orders of magnitude smaller than datacenters use today.""
Would this DNA also exist in the same conditions as the data center or does it need more things?
To put it into context, each cell in your body contains 6 billion base pairs of data (two copies of your genome). Each base is one of 4 bits, so that’s 4^6000000000 of data in each cell. Your body has ~37 trillion cells [1]. A person is about the size of a rack (well, maybe 10-15U by volume), so that’s 3.7e13 * 4^6e9 bits per rack.
A petabyte is 8e15 bits.
That’s a lot of data storage capacity in a small space. Moreover, there is the potential for introducing more synthetic bases to increase the 4 to 6.
With 6 gigabases per cell and 2bits per base, the storage capacity is 12 gigabits (1.5 gigabytes) per cell. And with 37 trillion cells (3.72e13) in a human body, that’s 3.72e13 cells * 1.5e9 bytes per cell which is 5.58e22 bytes per person or 55 zettabytes. This seems like a more reasonable number.
Of course even when you said base you did the math with # of base pairs, so it doesn't really matter.
T-A
A-T
G-C
C-G
That location, using 2 base's, encodes 4 possible values, i.e. 2 bits. I.e. 1 bit per base.
But that doesn't explain storage density. And biological DNA has mechanisms to survive in a cell but these synthetic DNA's might be standalone.
The meme is true
You’d have to deconvolute the signal first, then computationally determine the strand you actually read.
Doesnt really make sense to switch to a different copolymer.
There are really three steps: writing, storage, reading.
Writing: There are dreams of using biology to write DNA (currently we use chemical synthesis). See the San Diego startup Molecular Assemblies.
That could miniaturize a big part of it and not require your body to host bottles of nasty chemicals. Maybe you could find clever ways of siphoning the chemical building blocks and ATP energy from the host.
Storage and retrieval: maybe just have bioinert containers implanted. You’d like to be able to reuse them so ideally there would be a robust flushing mechanism.
Reading: the Minion reader from Oxford Nanopore that this paper uses is about the size of a usb stick. The technology inside is a hybrid of electronics and protein. This technology will undoubtedly get better and smaller as sequencing has a a lot of R and D money going into it.
I'm not being a luddite here on purpose, but over long time scales there's a tremendous potential for this kind of technology to push towards a kind of class differentiated society in the way most of us would despise.
Some technologies are leveling, like roads or mass transit or vaccines or industrially produced consumables. I don't see public institutions putting libraries in the seeds of apple trees as a civilizational fail safe, whether that's centrally planned economies or democracies. But maybe you could get an ethnostate like Israel to include the talmud in your cells or your microbiome when you settle on occupied land. Best case scenario with body horror is that it becomes like tattoos. I await the forthcoming Atwood book with that slightly alarmist slant.
Once humanity becomes capable of living and thriving in different planets/outer space without the mother planet is when things will start getting really interesting.