Harvard cracks DNA storage, crams 700 terabytes of data into a single gram
extremetech.com
extremetech.com
while that's true (re: scalability), it's not what we did, and would be hard to scale.
Clearly with some form of fountain code or LDPC codes you will be able to get the data back, but what struck me is that I always thought of DNA as relatively unstable, in the sense that cells decay/die etc, but the fact that just sitting there, DNA which isn't expressing various proteins under the influence of other cellular mechanisms, well it just sits there. That was new for me.
When I showed it to my wife she pointed out that the sourdough starter she has been using since we were married was from her grandmother, I joked that the next megaupload type raid would have to sequence all the DNA the found in a place to figure out if Shrek3 was encoded in it somewhere. That would be painfully funny I think.
If you needed more, you could then transform E. coli with the plasmid and let them do the work for you.
I don't know if this is done very much anymore, but this helps to show just how robust DNA can be. I want to say that a few years ago, someone tried to sell our lab an archival tool for DNA that was essentially this. You dried your DNA samples onto blotting paper in a grid, then when you needed a sample, you could go back and reconstitute by punching it out of the paper.
I don't know if there is a copy outside the Science Paywall, perhaps Sriram could answer that.
If one does need multiple copies, it would seem that this method suffers from the coupon collector's problem [1] (i.e. to collect all 50 strands requires collecting 225 random strands on average), and that the retrieval rate could be improved by using a fountain code [2], which allows each strand to simultaneously encode data at multiple addresses, which would decrease the number of strands required to be sampled to only slightly more than the number of strands worth of data requested.
[1] http://en.wikipedia.org/wiki/Coupon_collectors_problem [2] http://blog.notdot.net/2012/01/Damn-Cool-Algorithms-Fountain...
Holy shit, this exists? Wow.
More and more often I find myself not reading articles because someone thought it would be a great idea to create a non-scrolling, non-obvious, paginated "iPad format" with additional misleading and unintuitive buttons looking like native ones but doing something different.
Sorry for the rant.
EDIT: so you might as well access the original at http://hms.harvard.edu/content/writing-book-dna instead of the ad-ridden regurgitation.
paper http://db.tt/ZDoDJZeD supplement http://db.tt/elIqsy72
That's not to say that we weren't seeded, just that if we were, any message would likely have been lost. Unless... maybe they seeded mitochondria intact, which are in all (ok, most) eukaryotes. Maybe that might work... :)
Unfortunately, chrM is pretty small too... so no hidden messages there either :(
you are either the most humble senior author ever or jerking all our chains. possibly both. congratulations either way.
T = 00
G = 01
A = 10
C = 11
for example.AT = 00
TA = 01
CG = 10
GC = 11
The trick would be to always correctly identify which is the left and which is the right strand. I don't know if that is possible in practice though.
You just encode a big marker (making sure it's not a palindrome-paired version of itself!) as a header. If you see that, it's a correct order. If not, it's not.
[Left strand]
A = 00
T = 01
C = 10
G = 11
[Right strand]
T = 00
A = 01
G = 10
C = 11
Anyone know these guys at Harvard, b/c this might be a way to put, at most, 2800 terabytes in a gram? (I don't know how long the header sequences would have to be).
I can't find a copy of the original article, but it's certainly more informative than some science journalism fluff piece.
Let's say DNA breaks, or some errors appear in the code. Thanks to the double stranded structure it's "quite easy" to repair the code.
Besides that, it's not the density which is a problem right now, but the access speed. The amount of data in DNA is so immense that doubling the density won't give any practical improvements for decades to come - if ever.
Having said that, if I'm not mistaken, some viruses are encoded by single stranded DNA & ssRNA. I'm not sure, but the density might be the reason for that.
Assuming 5.5 petabits stored with 1 base pair representing 1 bit, we can extrapolate the time required to extract the data based off the time taken to sequence the human genome (3 billion base pairs).
5.5 petabits / 3 billion bits ~= 2 million, so theoretically it should take 2 million times longer to sequence the original.
3 years ago, there was an Ars Technica article about how it now only takes 1 month to sequence a human genome[1]; the article now claims that microfluidic chips can perform the same task in hours.
Assuming 2 hours (low end) to sequence the human genome:
2 hours * 2 million = 4 million hours = 456 years, give or take a few years.
So, maybe not so great for storing enormous amounts of data. But if you want to store 1 GB, it would only take ~6 hours. Not too bad.
[1]http://arstechnica.com/science/2009/08/human-genome-complete...
Even if it doesn't, you'd still need something like 5,000 experiments in parallel for it to take less than a month...
1) DNA sequencing technology is currently advancing much faster than silicon technology, so give it time and it's likely that it will catch up with current hard disk reading speed at comparable sizes and volumes. 2) DNA's self-hybridizing nature makes it easy to pull out blocks with specific addresses (if you wait for the hybridization). So if you include address labels in the DNA as you write it out, you can probably pull it out in chunks of kilobase to megabase at a time. 3) As the other commenter pointed out, this is extremely easy to parallelize. So if you want to go twice as fast, divide the sample in half put half in each machine. Dilute and pipette as necessary.
paper http://db.tt/ZDoDJZeD supplement http://db.tt/elIqsy72
tldr; we are 6-8 orders of magnitude away from doing petabytes routinely; that said costs of sequencing/synthesis have seen such drops over the last decade or so. there are many barriers though for that continuing for another decade
"In humans, the deoxyribonucleic acid (DNA, Germ. DNS) is the carrier of genetic information, and the main constituent of the chromosomes.
DNA is a chain-like polymer of nucleotides, which differ in their nitrogen bases (Thymin/Cytosin bzw. Adenin/Guanin,) The alphabet of the code is therefore: {Thymin, Cytosin, Adenin, Guanin,} or also { T, C, A, G } Three consecutive bases form a word So there are 43 = 64 combinations per word so the word length is ld (64) bits = 6 bits A gene contains about 200 words A chromosome contains about 104 to 105 genes The number of chromosomes per cell nucleus is 46 in humans The stored data per nucleus have, a volume of 6 bit * 200 * 10^5 * 46 = 55200 bit * 10^5 * 5 * 10^9 bit * * 10^9 Byte = 1 GByte"
Human genome is 23 chromosomes, 2 copies of each (46). They are not mirror copies like RAID1, but rather alleles (different versions of same gene). So you might have a different version of a gene on each half of the chromosome. This is the basis for sexual selection and why sexual organisms can evolve (esp at the population level) so much faster than asexual ones. The two alleles can be the same as well. It's called homozygous or heterozygous. This is basic Mendelian genetics. Most genes work this way, though many are more complex.
ATGC is correct, so each "bit" is base 4.
Genes encode proteins; every 3 base pairs is a codon which specifies which Amino Acid to use. While there could be 4^3=64 in practice there are only 20 amino acids used in nature to make functional proteins. Genes vary greatly in length, not sure where you got ~200 codons/gene but that's terribly wrong. Maybe close to average, but the range is large. In any case, for data storage, anything relating to codons and proteins would be irrelevant.
Also in practice not all of a chromosome encodes proteins. There is often lots of buffer region between genes, not to mention a lot of control flow sequences that help control expression of genes. Beyond that, the ends of chromosomes don't have many genes, they just contain pseudo-random information and are still being explored (junk dna, telomeres, etc).
Info: According to Quora the avg. weight of the Human DNA is 60g, that means a Human could carry 700TB * 60g = 42000 TB Note: 700TB = 7.69658139 × 10^14 bytes
But hard drives aren't the densest storage medium in use today. A microSD card can hold up to 64 gigabytes and is 0.5 grams. 700 terabytes would be only 5.6 kilograms.
Experience and creativity would be what everyone garners to get a job.
If it did, why, you'd never have to study protein sequences- you'd already know how to make hemoglobin!
Not saying it's practical or desirable, just that it is possible.
Edited to add: Ruined my own punchline. Steganographic stegosaurus, damnit!
I must not fear, fear is the mind killer...
3D vs 2D is also a good point, made by another in this page.
I'm gently concerned about what'll happen to information if it's not available to the future people. Is anyone taking the most important documents of our civilisation and encoding them onto clay tablets, or some such?
I imagine this could be solved by ensuring that every helix starts out with something like a byte order mark that would distinguish the two strands reliably.
Storage density could be further increased by a constant factor if more kinds of bases were used, or if they could use single-strand DNA/RNA instead (which would probably require some chemical means of ensuring that a free strand doesn't accidentally bind to something else).
How much is this news true and how much is it the usual ExtremeTech editorialism?
For instance does DNA really last forever?
The progress in this field has been far better than Moore's Law [1] so it could be very practical.
[1] Excellent TED talk on it too by Richard Resnick: http://www.youtube.com/watch?v=u8bsCiq6hvM
Non-snarky reply: There is a huge difference between phenotype stability and genotype stability. Take 100 cells from your body and you'll find hundreds if not thousands of genetic differences between them (single mutations).
You have to remember that offspring are only going to have the best DNA passed down to them, since many mutations would result in non-functioning gametes or non-viable offspring.
The best example I can think of spermatozoa production. DNA is copied in that process and a large % of spermatozoa are non-functional.
(How nobody has reference this?)