Harvard stores 700 terabytes of data into a single gram of DNA (2012)
extremetech.com
extremetech.com
Much like figuring out how much to tip your waiter is a mathematical exercise. Just one that might not shake the foundations of academic mathematics.
Just because territory is well travel, doesn't mean everyone has walked there.
"Wikipedia trivia: if you take any article, click on the first link in the article text not in parentheses or italics, and then repeat, you will eventually end up at 'Philosophy'."
I know a more then one prof who'd agree with that statement.
OK, so assuming 1 hour for 3 billion bases, it's 1000 hours for 3 trillions bases (3 Terabits) or 1 million hours for ~3 petabits (or 400TBs of data). Yeah, that's a long time, roughly 100 years :)
Seeking can be performed through entirely different means, as DNA is content-addressable. You can put in a magnetic bead attached to a strand complementary to what you're seeking, and pull that out of the mix. This is still a physical process that can take quite a bit of time.
Let's be clear: this work encoded just 5.27 megabits. It's stored in what's basically a large molecular hash table where each piece of key-value data is replicated a million times for redundancy. Each piece then read 100x to correct for the -abundant- errors in each piece. So they encoded less than a megabyte.
The problem with encoding information into DNA is that writing serial polymers accurately is difficult and slow. In this paper they're using an inkjet printed DNA array. It takes a day to make them, resulting in a bandwidth of:
(5.27 Mbits) / (24x60x60 sec) = 66 bits/sec
Reading is a little faster. The fastest system, the HiSeq 2500 reads 120Gbits raw in 27 hours. Factoring in the necessary 100x redundancy, one has a -maximum- read rate of:
(120 Gbits / 100) / (27x60x60 sec) = 12 Kbits / sec
So for 5.5 petabits it would take 16,000 years to write the data into a cubic mm but only 16 years to read it at the current rate.
If we get a little scifi, and assume we build programmable polymerases and get nanopore (direct read) sequencing. Even then, physics limits you to something like 1000 read/writes per sec per pore/polymerase. Instrumenting to these will probably limit per-feature size to being larger than 100micron on a fabricated chip, giving us an ultimate read/write limit around:
(2cm/100um)^2 x 1000 bit/sec = 40Mbit/sec
for a giant 2cmx2cm chip. With the necessary error-correction and redundancy, it's probably going to cap out around 1Mbit/sec at best.
Those 700TB take about 2months to read/write at these rates, and we'll have much-better solid state storage technologies by the time we figure out how to all that with DNA.
X digits of binary vs quaternary can represent 2^x vs 4^x possibilities.
Even at 10 digits, that is 2^10 = 1024 vs 4^10 = 1048576.
with base2 we get 2^5 = 32 possible combinations
with base4 we get 4^5 = 1024 possible combinations
2*(2^x) != (4^x)
Take, for example, 32bit integers vs 64bit integers (unsigned for simplicity). Two 32 bit integers can represent exactly the same number of combinations that a single 64 bit number can. Sure, there's an exponential number more combinations in a 64bit integer than a 32 bit, but the number of combinations is not how the storage capacity is measured.
EDIT: Yes well for those who haven't read the book the reference is to the sexual transmission of data. Implied (but not confirmed) by nano-tech, however no implementation details were provided (as I recall), so encoded DNA might have been a possibility as a medium. Just sayin'.. /shrugs/
[1] http://stackoverflow.com/questions/8954571/how-much-memory-w...
This is pretty interesting. I'd love to read the paper but it looks like I can't download the article yet since its not finished uploading to the site.
What kind of IO bandwidth can you currently get from DNA? Does reading from it damage the DNA? What kind of hurdles need to be overcome to bring this to market and are they hurdles we can overcome in the near future?
edit; Hey looks like you answered most of these else where, so thanks.
Of course, it's all about the energy. I'm not a biologist, so I won't speak to the details of how resilient DNA is to specific types of exposure or why, but I am a physicist and I can tell you that if you blast DNA with high energy radiation, like gamma rays, it's gonna have a bad time. That said, if we want to talk about robustness in terms of what it's likely to be exposed to, then it's pretty darn robust. We are exposed to a wide band of EM radiation on a constant basis, and can even withstand exposures on the higher end without becoming 'corrupted'.
I'd also note that standard, non-specially designed electronics aren't very resilient to radiation either. Not to mention, if you could reliably read/write data using DNA, I'd expect that redundancy would become... easy. Security on the other hand...
Sneeze once and everyone has a copy? :)
This probably required a machine significantly larger than most desktop computers, with a pricetag in the $100,000+ range. That and it has a R/W speed of ~30Kb/s...
It will be in our hard drives when the technology has been developed to make that happen.
I mean, graphene is already looking to be a far more efficient and powerful alternative to silicon. Why isn't that in our processors yet? It's because creating a logic gate and creating an x86 CPU are entirely different things.
The same way that encoding a bunch of data into DNA and replicating it a few billion times is entirely different from making an on-demand biological data storage device.
Any idea on how long this will take to get to market, and whether it's being worked on?
Storage companies buy by the petabyte.
The argument that we don't need more storage doesn't work.
I said that all-you'll-ever-need capacity is meaningless if you can't get what little you need without extreme efforts. It's a "drinking from a firehose" problem.