500 Exabytes per Raindrop
math.ucr.edu
math.ucr.edu
I was talking to some of our systems engineers about the size of the Hadoops clusters at Yahoo. I was pretty impressed with having access to thousands of machines with 10s of petabytes of storage (think kid in a candy shop).
It's nice to be put gently back into my place, by a simple raindrop.
Modern instruments are ever more computerized and spit out gigs and gigs of 1s and 0s with every use.
And then you have to take all that data and try to turn into something humans can understand.
And every time our instruments get better the amount of data we collect goes up.
Systems biology is basically a data management problem.
Just imagine how much bigger our storage mechanisms will get...
Isn't that what the author was getting at?
In fact, it's possible to say then that you could start with something as simple as, say, a piece of fairy cake, and deduce accurately the nature of everything else in the universe.
Reading this, the thing that stuck out to me was the relatively small size of text (5MB for Shakespeare versus 20GB for Beethoven). It struck me that poetry - particularly haiku, which deals with nature - is sort of a primitive way of applying lenses to the world to create a filter for viewing things by appealing to collective experiences. Just a thought.
Don't get me wrong: poetry has become incredibly sophisticated. And it's still as enjoyable to write as it was long, long ago.
"I'll probably keep working on topological quantum field theory and other wimpy subjects"
I would be careful what you tell him to go read...
And a heaven in a wild flower,
Hold infinity in the palm of your hand,
And eternity in an hour
William Blake, Auguries of Innocence
I put "information" in quotes above because the actual "information per bit" of English is pretty low, and that's where the ability to both compress it and comprehend it, not as single bits but as groups of bits, comes from. Still, compressed data is still comprehensible once you uncompress it, so it's a measure of the surface comprehensiblity. It's really a measure of information density. Compressed data has a high information to space ratio, whereas random, uncompressible data has low information to space ratio (or perhaps negative).
You can experiment with this yourself:
cat > /tmp/string.a
(paste in some English text cut and pasted from a web page)
dd if=/dev/urandom of=/tmp/string.b bs=1 count=`wc -c < /tmp/string.a`
bzip2 /tmp/string.*
ls -l /tmp/string.*
Experiment with the above for corpora of different lengths. You'll see that as there is more English text, which is information rich, it can compress better than shorter English text (as a compression ratio), and will compress better than random data (which we know contains little information) of the same length.A string composed solely of 27 As would compress down to perhaps 2 bytes or less (not including the size of the decompressor). You are right: there is not much information in it. Less than 16 bits of information in 27 As.
http://en.wikipedia.org/wiki/Information#Measuring_informati...
If I look at the occurrence of subsets of the string, then that would be a better discriminator: the random string's subsets should follow a normal distribution while the English string's subsets will be highly skewed.
However, that doesn't work when I try to discriminate between an English string and a string generated by a simple algorithm, since the latter's subset distribution will also be highly skewed. What kind of metric discriminates the English sentence from either case?
Am I making sense here? I haven't had any formal training in information theory, and my brain is kind of fried right now.
A raindrop is made up of a certain number of atoms, but these atoms are all the same, and cannot store any information. To store information, you need items that are dissimilar to each other. So the comparisons he makes are not correct.
If raindrops were perfect crystals at absolute zero (and there were only one isotope each of hydrogen and oxygen), they would be a lot less information in them.
There is a very small amount of information in a raindrop, no matter how complex it is from the molecular structure.
The entire argument is flawed. It's based on a completely wrong premise - information is not the same thing as structure.
I can't explain this any better, you will need a leap of intuition to understand what I mean, but when you get it, it will be obvious.
Ultimately, one should come away with the understanding that obviously not all that information is important.
6 years ago, oh my.