A cheap way of storing information in DNA
qz.com
qz.com
“What Catalog does, instead, is cheaply generate large quantities of just a few different DNA molecules, none longer than 30 base pairs. Then it uses billions of enzymatic reactions to encode information into the recombination patterns of those prefab bits of DNA. Instead of mapping one bit to one base pair, bits are arranged in multidimensional matrices, and sets of molecules represent their locations in each matrix.”
"Using nanostructured glass, scientists from the University’s Optoelectronics Research Centre (ORC) have developed the recording and retrieval processes of five dimensional (5D) digital data by femtosecond laser writing.
The storage allows unprecedented properties including 360 TB/disc data capacity, thermal stability up to 1,000°C and virtually unlimited lifetime at room temperature (13.8 billion years at 190°C ) opening a new era of eternal data archiving. As a very stable and safe form of portable memory, the technology could be highly useful for organisations with big archives, such as national archives, museums and libraries, to preserve their information and records."
IBM has also demonstrated manipulating individual copper atoms on a surface and arranging them into patterns - a scale that would be much denser than organic molecules.
I have seen electron microscopy setups 10 years where you could cut individual lines on a microchip with ion-beam milling and "draw" another one with 7nm with via Gallium deposition (if I remember right). All that is impressive technology with specific applications but not good for scale up.
When creating ustom DNA because a less a cumbersome process, it could be a cheap and space-effective way to do cold storage (both literally and figuratively).
Also, it'd be a quite sneaky way to transfer data. Spray some DNA on someone's jacket, go through all the security measures and afterwards, swab the jacket, clone and sequence and bob's your uncle.
However the basic density isn’t really the factor of interest. What’s more interesting about DNA is that you don’t need to store it on a surface. So while with semiconductors, you can only really cover a 2d surface (and maybe stack a few dies for density). With DNA you can pack a lot more information into a 3 dimensional area.
I however am not particularly bullish in DNA storage. We don’t really have the read systems we’d need to make this viable (still costs ~1000USD to read the ~ 1Gigabyte human genome and > 1 day). Unless these guys have really developed something very novel, we don’t have the write systems either.
Both the read and write systems are also highly errored compared to digital storage (like error rates of 1 in 100). So their are many issues...
This piece is dropping fast though; Dante Labs will currently do it (with 30X coverage) for individuals for $500 [1]. And I believe the price for labs that do a lot of sequencing is significantly lower per-genome.
Writing is still a much bigger challenge, though. Writing something the size of the human genome is currently considered a "grand challenge" and is being tackled by HGP-Write:
https://en.wikipedia.org/wiki/Genome_Project-Write#Human_Gen...
[1] https://us.dantelabs.com/products/whole-genome-sequencing-wg...
Illumina appear to be targeting 100USD right now. I think their markup on reagents is probably at least x10...
But even that’s super expensive compared to reading data from a HD... so you have to wonder where DNA as storage is viable. Unless a vastly cheaper read method becomes available.
Yes compared to HDD it's way too expensive to read. However tape backups are also cumbersome and time-consuming to use, but the fact that you can get extremely cheap bulk storage from them means lots of big companies still use them. So I could see it maybe being competitive there (eventually).
"but how do you prevent mutations" Redundancy. Remember, a mol contains quite a few molecules. a nmol (nanomol) still contains a few. Even an amol or whatever contains a decent amount :-)
A recent approach (https://www.biorxiv.org/content/early/2018/06/16/348987) includes synchronization nucleotides. This constrains the sequence space, aiding recovery of data from multiple strands, even if those strands contain errors (provided that they do not all include the same error).
Additionally, different types of errors are more likely in certain contexts, depending on the synthesis and storage approach. This particular paper utilizes enzymatic synthesis (which is the new hotness in DNA technology) and the authors model these error probabilities.
Combined, with a couple of other nifty tricks, they demonstrate ~30% error tolerance across 10 strand variants (i.e. copies of the same data containing errors), albeit for a very short sequence.
> If the DNA isn't being biologically replicated, why use DNA over any other 4 distinct chemical structures? In many cases, the DNA is being biologically replicated -- not within organisms, but biochemically via enzymatic reactions. One attractive feature of DNA storage aside from its density is that it's cheap to make a ton of copies. After all, Nature does this all the time as cells divide. And given the cost of mutations (most are detrimental), DNA polymerases have evolved (ironically?) to be high-fidelity. Something like 1 error in 10^8 base pairs.
And that touches upon a broader point: DNA is an attractive substrate for alternative storage relative to other polymers because its so widely utilized in nature and, as a direct result, we've developed a lot of technology and infrastructure around it. However, much of this development has been around sequencing (reading), alignment and assembly algorithms, and, though still sorely lacking, biological interpretation. It's exciting to see growing interest over the past few years in DNA encoding and synthesis. Remains to be determined which approach(es) become workable standards for the field.
Here's the requisite quote from Echopraxia, by Peter Watts:
> Even DNA computers, custom-built for a specific task and then tramped carelessly into wild genotypes like muddy footprints on a pristine floor. Nowadays it seemed like half the technical data on the planet were being stored genetically. Try sequencing a lung fluke and it was even money whether the base pairs you read would code for protein or the technical specs on the Denver sewer system.
With typically three base pairs per amino acid, encoding complex proteins in 30 base pairs isn't possible. That said there are other methods to assemble longer sequences.
Nothing is technically stopping someone from writing a deadly virus, except that you'd have to engineer its DNA (not a trivial task). Why do that when nature has created plenty of deadly viruses to choose from? Furthermore, for military purposes, you would want something controllable, so you can deploy it without also killing yourself. Bacterial agents tend to be more suited to this.
It is actually trivial. You can look up the DNA for most viruses. If it is long you could just ligate a few oligios together. I wonder if oligo synthesis companies have black lists of specific sequences if a random person orders them.
Also, many viruses are RNA based. More tricky than DNA. But in that case I would just order DNA, build a complementary DNA virus and then use RNA polymerase to make the RNA.
" for military purposes, you would want something controllable"
I am not sure the guys that I am thinking of that would be interested in such a thing would worry too much about that. I guess if you are willing to fly a plane into a tower then you would not be too concerned of dying in the process of developing a virus.
Better than a virus may be the Mutagenic Chain Reaction. Maybe my genes are better than yours and I want to spread mine? ;-)
http://science.sciencemag.org/content/348/6233/442
"The threshold necessary for small groups to conduct global warfare has finally been breached, and we are only starting to feel its effects. Over time, in as little as perhaps twenty years and as the leverage of technology increases, this threshold will finally reach its culmination -- with the ability of one man to declare war on the world and win." Brave New War back in 2006
I had assumed here that "writing" meant engineering new code, not copying from existing viruses.
Terrorists getting hold of viruses has always been possible. If someone were so inclined, I'm sure they could get hold of an Ebola strain and pass it around somewhat, using no special technology.