A biological camera that captures and stores images directly into DNA
nature.com
nature.com
Obviously if you have a special use case, then that may dominate your radix economy (like hex, b64, etc...), but for general purpose information purposes, the order base 3, base 4, then base 2.
This present a lot of interesting questions to me. Like, why didn't DNA end up as base 3? (probably because 4 naturally lends itself to pairs of 2).
Also, this idea of radix economy goes beyond just the encoding of information and is represented in logical economy as well. So for example, ternary logic is (much) more efficient than binary logic. Having that 3rd state just makes problem solving much more elegant.
To that end, I have always wondered how nature has exploited this 4-state number system logically. Like, are there all sorts of exotic logic gates that come from a 4 state system?
By using base-4, there's enough space to permit lossiness of the coding itself - given the number of amino acids and the 3-NT encoding.
So you really aren't optimizing JUST for nucleotide encoding, but you're also optimizing in concert with 3-nt/AA, and 20AA codes.
So if you have to optimize for information density and fidelity, given X-nucleotides, Y nucleotides/AA, and Z AAs, and sample as much chemical and physical diversity in those AAs life has settled upon: X=4, Y=3, Z=20.
If we went with X=3, you might need Y=4 to get the same kind of fidelity, but that cranks up your energy costs by 30% (from 3 to 4 NT per AA).
Apart from repairing structural damage such as missing bonds, the cell can even repair missing bases or non-straight breaks without loss. This mechanism is also used for replication: the entire strand is split and each half is completed with its mirror counterpart.
Some DNA is used to attract other proteins, or even interact with DNA elsewhere on the strand, or is translated to RNA (one-on-one) which can then have a function based on its sequence or the structure it folds into.
Any 'logic' there is, is built _on top_ of this.
Why did we end up with only 20 proteinogenic amino acids? Why are vertebrate neural architectures inverted (cell bodies on the inside, connections on the outside, even though the other way round way (eg. like a squids brain is organised) is easier and less inhibitive to growth?
2 Reasons:
a) Because nature and evolution cannot engineer. Random mutation, recombination and natural selection are the only mechanisms available. Things get selected if they outcompete existing alternatives, they don't need to be the best solutions.
b) All solutions have to be built by modifying what already exists. Evolution doesn't get to do greenfield projects, because anything that has to start from scratch is so disadvantaged in natural selection compared to already evolved complex life, it will fail.
This leads to systems that, from an engineering point of view, don't always make a lot of sense.
Eg. the architecture of the vertebrate neural system creates a lot of issues (eg. our light sensitive cells point in the wrong direction). The only way this makes any sense if when one looks at how the neural tube (the precursor to the backbone) is formed by the endodermis folding in on itself. This process is so deeply at the root of the Chordata, and so many other things depend on it, that it simply cannot change any more.
Many many biological systems are "legacy systems" in the truest sense of the word: Solutions produced a long time ago that may have many problems, but are simply too deeply enmeshed with everything that came after, that they are now impossible to change.
Mind elaborating on that?
Because there is no biochemical reason why DNA could not have incorporated, say, a third pairing pair, so while base-3 (which I don't specifically mention in my post btw.) wouldn't work, base 6 or 8 would have been possible. "Unnatural Base Pairs" are even known to work in laboratory settings.
There is also no biochemical reason why base2 life wouldn't work. Expand the reading frame of the translation machinery to 5 instead of three, and you have enough coding space for polypeptides.
My answer adresses the question completely, because the only reason behind these "decisions" is an ancient system that simply got "frozen", and now cannot change any more.
are you sure about that? are you sure there's no weird effects that might destabilize very long sequences of 2-nucleotide DNA? or on how wide DNA-binding domains have to be to cope with reduced information density, and how that might sterically hinder smaller arrangements of proteins?
> My answer adresses the question completely, because the only reason behind these "decisions" is an ancient system that simply got "frozen", and now cannot change any more.
your answer is just a hypothesis, not a proof. these things can be studied (by studying abiogenesis in-vitro), and it's not certain these decisions were "flash frozen" like you describe. 2-, 4-, and 6- nucleotide coding systems might have coexisted in the RNA world, and 4- could have won out for some reason.
Yes, I am sure about that, because I used to study Biology before going into IT. And we had a lovely lecture in which we used to discuss theoretical setups for lifeforms at a molecular level.
2 nucleotide DNA isn't necessarily less stable. AT-rich domains have less bindings, but if stablity is the issue, use CG instead (3 bindings)...although that is also a compromise, because then opening DNA for transcription gets more difficult.
> your answer is just a hypothesis, not a proof.
My answer is what we observe in evolutionary biology.
I have given an example outside of the molecular world for a reason. There is no real advantage to the inversion of the neural architecture in Chordata, it just didn't matter when the neural tube formation mechanisms came to be. Now, with mammals having huge brains and complex sensory organs, the warts in that design show.
The proof for that is easy to come by, (also a reason btw. why the neural inversion is my favorite example for this): Look an any Protostomia. Their neural system isn't inverted. Consequently, Squids don't have a visual blind spot.
for instance, I've seen arguments that the codon mapping, and even the particular set of protein- coding amino acids, that we ended up with was arbitrary, but I've also read papers arguing that the amino acids include a sort of spanning set of different structural scaffolds with different polarity that happen to mesh well with DNA, and that the particular choices of codons were influenced by how the RNA t-acyl transferases arose, etc.
so, I'm still unconvinced, but I find this area fascinating to read about.
> your answer is just a hypothesis, not a proof. these things can be studied (by studying abiogenesis in-vitro), and it's not certain these decisions were "flash frozen" like you describe. 2-, 4-, and 6- nucleotide coding systems might have coexisted in the RNA world, and 4- could have won out for some reason.
His hypothesis is, at least in part, “4- won for some reasons for which we have no explanation, and it stayed that way for some reason [that we may or may not know].” I suppose the reason would be that 4- was somehow better suited for the particular use-case at the time.
Of course there’s a ton of interesting details to discuss to discover, and whether if multiple systems coexisted is one of many fascinating things to discuss, and his response never said otherwise.
That's true, but a) not the point I am making, and b) I am pretty sure it says nowhere in my post that it would.
Can you expand on that? Are you talking about front-facing eyes vs. birds' eyes? Or something else like retinal structure?
https://en.wikipedia.org/wiki/Retina#Inverted_versus_non-inv...
https://en.wikipedia.org/wiki/Retina#/media/File:Retina-diag...
Edit: ninja'd
But that doesn't mean the setup makes sense, and that is exactly my point.
And long term, this has an impact. For example, vertebrate brain size is limited by the simple factor, that we have to put all the connections on the outside. The more neuronal bodies we have, the more connections they require.
N <---> N
In this clumsy diagram, 2 neurons talk with the connections on the inside. However, vertebrate brains have to do this instead: +--------------+
| |
+-> N N <-+
It's easy to see how the second setup becomes prohibitive when more Neurons are added to it. The brains of Protostomia again don't have that problem...they can have the connections on the inside, and the neuron bodies on the outside, aka. the logical setup.Now there are ways around that, eg. Reptile and Bird brains grow in bulbs that theoretically allow sustained growth without the connective layer getting in the way. But similar to the reflective layer in our eyes, this is not a setup that's there because it makes a lot of sense...it's a hack, a workaround for some "legacy system", that is now so enmeshed, it's impossible to change.
Lefthand is a vertebrate eye, righthand is a squids eye.
In Vertebrates (really in all Chordata), the light sensitive "tips" of the sensory cells point inwards, aka. the exact wrong direction. At the base of the cells are the axons (nerve connections) which transmit the information into the brain.
Due to the aforementioned orientation, these axons run along the outer layer of our light sensitive cells, and at some point have to travel "invards" towards the brain. At that point there can be no cell bodies, and that's the "visual blind spot" of our eyes.
A squids eye doesn't have that problem; all the light sensitive cells point outwards, the axons are at the innermost layer, and connectivity can be achieved without a blind spot (also, they don't need a reflective layer).
Hubris and attempts to alter inherent nature are often tied up ironically. But we can benefit a lot more from biological humility, realizing there are many unknown unknowns.
One, is that base 4 makes a lot of sense for the stability of DNA structures. You have two purines, two pyrimadines.
Another is that partly because codons are degenerate, the distribution is way off a uniform distribution. For chemistry and mol bio reasons, the distribution of AGTC is very skewed.
When i fully wake up, this might be a fun blog post to draft.
Which is a problem given DNA is a lossy format.
This ratio is what leads to base e being the theoretical maximum.
https://en.wikipedia.org/wiki/Repeated_sequence_(DNA)
Many real wire protocols have mechanisms to prevent repeated sequences entirely
https://en.wikipedia.org/wiki/8b/10b_encoding
DNA coding for real proteins is unlikely to be too terribly repetitive but I image a long α helix could have a repetitive amino acid sequence. Many amino acids can be coded with variant codons, I guess if repetition were a problem in a particular gene natural selection could step in.
https://en.m.wikipedia.org/wiki/Trinucleotide_repeat_expansi...
TLDR: DNA triplet repeat expansions cause diseases like Huntington's disease.
Binary for electronics is obvious because there are 2 states in electric components: on or off. There is no 3rd option.
DNA base set, Amino acid set, Translation layer between DNA/Proteins.
Currently, we've got: 4 DNA bases, 3 bases/AA, 20 AAs; 4^3 => 20
If you change one of those numbers, you'll need to rejigger the rest, and you'd need to reoptimize. And there are competing goals which at least include: - maximize access to biophysical/chemical diversity - minimize energy expenditure to produce each component, chemically - minimize energy expenditure to both copy instructions & produce products - maximize information fidelity - minimize or at least degrade gracefully in the context of errors
In the context of a 3-base system, you very well could throw off those optimizations given the consequences for the other 2 parameters (#AA & nt/AA). 3^3 = 27, which is very close to the maximum of 20 amino acids. Which means you'd probably need a 4nt->AA translation layer to keep the same number of AAs, and that alone would add 30% more energy expenditure. If you kept the 3nt->AA system you'd BOTH need to reduce the number of accessible amino acids AND you'd lose some of the error correction mechanisms of having degenerate codons code for the same amino acid.
We are all a bunch of biological goop resulting from random processes. Don't expect optimal solutions from evolution. There is no "why".
Ironic to say it’s not optimal, really we don’t have the full knowledge. Often when we learn more we learn how little we really know about biology.
Why would pairs of two be favorable?
> So for example, ternary logic is (much) more efficient than binary logic. Having that 3rd state just makes problem solving much more elegant.
What do you have in mind here?
> Like, are there all sorts of exotic logic gates that come from a 4 state system?
I don't know but you may be interested in this [1].
Seriously tho, using DNA as an information storage medium is a pretty neat concept.
And billions of years old.
With not quite good backup strategies.
This is done in a pretty obvious way, each "pixel" is a well in a 96-well e.g. plate and you expose the bacteria in these wells to different light and then the DNA transformation is triggered by the light, then you harvest the DNA from the bacteria and get your image pool library.
Except when it recombines weirdly and gets mixed up as per the article.
/s, but only slightly.
I mean, that would make your storage medium a potential biohazard. Although it probably would all be cool until someone put smallpox.bin on a major torrent tracker.
If it’s really that good, we should come up with a variant using slightly different chemistry so that biocontamination is not a factor.
generally no. If you're worried about random DNA being a biohazard there are way worse things to worry about, like how your immune system uses random stretches of biologically primed dna to create antibody diversity.
The real reasons why it's terrible is that write speed is atrocious and read speed is bad (on the order of 2-3x that of amazon glacier's robotic tape handlers, with WAY MORE expensive robots, and way more expensive cost to read -- you're bulk polluting rivers in china to make the reagents).
The only use case I can think of is deep generational archival (like the svalbard seed bank, but for information). Where cost to store by volume is at a premium, and where you'd like to have many many many copies, and you don't mind the cost to read because you won't be reading it but for every 10 or so years, if even.
Store your logs in DNA. You're never going to read them anyways.
100% agreed. I thought that was obvious, was mostly sniping at "DNA based digitalstorage startups", thanks for clarifying for me
I mean, DNA storage means you have an advanced, high performance DNA synthesis, copying, and sequencing machine. That’s fine in general, because most people using such a machine have a concept of the risks and responsibilities.
But imagine what would happen if we had millions of those machines connected to the internet with unpatched vulnerabilities. It seems like that could have some negative outcomes.
Padme meme with checksum
Yet while this fictional form is unlikely we have quite a lot of good examples and evidence for “inherited information”. You have to be careful with it since it’s too easy to accidentally include side channels for organisms to learn the information and thus break the test. Such as insects being genetically driven towards food by smell at a molecular chemical interaction level, and the smell becoming associated with the information you wish to test. A bee colony can’t be reliably tested unless you raise it from a new queen in an odourless environment if you wish to see if bees genetically know that the shape of a flower is associated with food. It’s tough to subtract the potential that a colony will have learned and “programmed” later generations of bees with things like the classic waggle dancing in order to more efficiently gather food.
We do have good ones though like cats and snake shaped objects, it’s surprisingly consistent, and pops up in some other animal species. It’s wired into our brains a bit to watch out for such threats. There’s a significant bias towards pareidolia in human brains and it’s telling how deeply wired we have some of these things, but it is there and study shows it seems to form well before our cognitive abilities do… these all have some obvious reproductive advantages however so it makes sense that the “instinct” would be preserved over generations as it confers an advantage. But it’s still impressive that it can encode moderately complex information like “looks like the face of my species” or “cylindrical looking objects on the ground might be dangerous”… even if it’s encoded in a lossy subconscious instinctual level.
I think it helps that the encoding does not have to be transferable in any way. This kind of "memory" has no need for portability between individuals or species - it doesn't even need to be factored out as a thing in any meaningful sense. I.e. we may not be able to isolate where exactly the "snake-shaped object" bit of instinct is stored, and even if we could, copy-pasting it from a cat to a dog wouldn't likely lead the (offspring of the) latter to develop the same instinct. The instinct encoding has to only ever be compatible with one's direct offspring, which is a nearly-identical copy, and so the encoding can be optimized down to some minimum tweaks - instructions that wouldn't work in another species, or even if copy-pasted down couple generations of one's offspring.
(In a way, it's similar to natural language, which rapidly (but not instantly) loses meaning with distance, both spatial/social and temporal.)
In discussing this topic, one has to also remember the insight from "Reflections on Trusting Trust" - the data/behavior you're looking for may not even be in the source code. DNA, after all, isn't universal, abstract descriptor of life. It's code executed by a complex machine that, as part of its function, copies itself along with the code. There is lots of "hidden" information capacity in organisms' reproduction machinery, being silently passed on and subject to evolutionary pressures as much as DNA itself is.
It definitely would be minimised data transfer, be it via an epigenetic nudge that just happens to work by sheer dumb luck because of some other existing mechanism or a sophisticated DNA driven growth of some very specific part of the mammalian connectome that we do not yet understand because we've barely got the full connectome maps of worms and insects, mammals are a mile away at the moment... no matter the mechanism evolution will have optimised it pretty heavily for simply information robustness reasons, fragile genetic/reproductive information transfer mistakes that work, break and get optimised out in favour of the more robust ones that don't break and more reliably pass on their advantage.
It can be useful to work from the evidence to a conclusion instead of the other way round.
But wondering and philosophising can be fun :]
It would be cool if humans could pass knowledge via their offspring. But I always get worried thinking if I'm the asshole, I wouldn't want my kid to be one too.
Plenty of mutations have no purpose whatsoever and were unrelated to survival or manifest after reproduction so are not selected for or against.
If you were, you would.
Recently here on HN someone posted a quote saying something like “if you shine light at something for a long enough time, don’t be surprised if you end up getting a plant”
It was about how the environment seems to reorganize in certain ways to use up energy (the latest Veritasium video about entropy also talks about this)
(I know viruses are RNA)
Yes, one of many.
Another one is a simple question: What exactly is the use case again? Because, storage isn't something we lack. Especially when talking about storage where, obviously, fast random access isn't a requirement, aka. data archiving.
We have good solutions for that; an LTO-9 tape can hold 18TiB of data native and up to 45 TiB of data compressed, with denser capacities planned: https://en.wikipedia.org/wiki/Linear_Tape-Open
"one gram of dried DNA can store 455 exabytes of data"
Seems like a pretty sweet use case to me!
I definitely do lack storage by the way. Say I want to download the common crawl data set, 380 TiB. And for redundancy I'd need multiple copies of the data too. That's a lot of disks for in the home. "18TiB ought to be enough for everyone" really doensn't cut it.
Yes, and half a gram of Hydrogen could produce ~500 Megawatts of power in a fusion reactor. However, that theoretical value will remain irrelevant, as long as we cannot build a practically useful fusion reactor. And even if we could build one, it still has to compete with all other forms of producing power for scalability, reliability, efficiency and cost.
The fact that there is a very high theoretical number that seems really impressive, isn't a use case.
So, with that being said: how long does it take to write these 455EiB? How long does it take to read them? How error prone are both processes? And how much does it cost to write/read them?
> "18TiB ought to be enough for everyone" really doensn't cut it.
Pretty sure I never said that.
Also pretty sure common crawl can be compressed. Even assuming only a 2:1 compression rate, that means it fits comfortably on 11 LTO-9's. Now, a quick google-search churned out tape prices of about 110-140 $ per LTO-9. Let's say ~150$ per tape, that means the whole thing fit's on 1650 $ worth of storage. About 5000 bucks with 2 backups included. Double that for uncompressed storage.
Alright, so how does that compare to DNA storage?
https://www.nanalyze.com/2023/03/dna-data-storage-solution/
quote:
These days, it costs $600 to sequence a complete genome which contains around 200 gigabytes of data or about $3 per gig. Today, magnetic tape technology offers the lowest purchase price of raw storage capacity at around two cents per gigabyte
end quote.
So just reading the 380 TiB back from uncompressed storage ONCE, would cost ~1,140,000 dollars.
And that's just for reading. At a price differential that is measured in multiple orders of magnitude, a technology better offer some REALLY good, REALLY tangible advantages to compete.
But an argument whether DNA is a viable option in the future or not would have to say technically what the issue of DNA is with future tech.
Whether it's more expensive today, or that there's no need for more data today, are not really arguments against it.
I do not intend to be arguing for snake oil or anything here though. If "DNA storage" is in a similar category of "perpetual motion machines" and "cars that run on tap water" then count me out.
I don't even know how our comments ended up being like arguing against each other. The only thing really I didn't agree with in the original comment was "Because, storage isn't something we lack", because I do find it lacking, both at home and in the cloud.
>The fossil record of Xiphosura goes back over 440 million years to the Ordovician period, with the oldest representatives of the modern family Limulidae dating to approximately 250 million years ago during the Early Triassic. As such, the extant forms have been described as "living fossils".[9] https://en.wikipedia.org/wiki/Horseshoe_crab
The inserted wiki-DNA could also conceivably be constructed so as to ensure its own perpetuation.
The goal of replacing memory cards is dumb, the tech that enables the storage is a foundational step forward in bio engineering.
Then they realise the picture is this dude: https://blogs.loc.gov/loc/2022/07/robert-cornelius-and-the-f...
This is the future. I don't think it will look exactly like this, and I don't think it will be here any time soon, but I'm excited to see these advancements.
What Cache is doing presently is trying to do archival storage in DNA - it has a lot of potential to be cheaper, more energy efficient, and more redundant. But some of the processes still aren't there yet.
Like we have the sequence of numbers for a jpg, but we've never seen the picture.