174355 pg11.txt
60907 pg11.txt.gz-9
58590 pg11.txt.zstd-9
54164 pg11.txt.xz-9
25360 [from blog post] 174355 pg11.txt
60907 pg11.txt.gz-9
58590 pg11.txt.zstd-9
54164 pg11.txt.xz-9
25360 [from blog post]You can convert any predictor into a lossless compressor by feeding the output probabilities into an entropy coding algorithm. LLMs can get compression ratios as high as 95% (0.4 bits per character) on english text.
There's no sense, for example, in which deriving a prediction about the nature of reality from a novel scientific theory is 'compression'
eg., suppose we didn't know a planet existed, and we looked at orbital data. There's no sense in which compressing that data would indicate another planet existed.
It's a great source of confusion that people think AI/ML systems are 'predicting' novel distributions of observations (science), vs., novel observations of the same distribution (statistics).
It should be more obvious that the latter is just compression, since it's just taking a known distribution of data and replacing it with a derivative optimal value.
Science predicts novel distributions based on theories, ie., it says the world is other than we previously supposed.
A statistical model of orbits, without a theory of gravity, is less compressed when you assume more objects. Take all the apparent positions of objects in the sky, {(object, x1, x2, t),...}. Find a statistical model of each point at t+1, so y = (o, x1, x2, t+1). There is no sense in which you're deriving a new object in the sky from this statistical model -- it is only a compression of observable orbits.
When you say, "if you have the new planet", you're changing the data generating process (theory) to produce a new distribution of points {(o' x1', x2', t'), ...} to include an unseen object. You're then comparing two data generating models (two theories) for their simplicity. You're not comparing the associative models.
Call the prior theory 8-planets, so 8P generates x1,x2,t; and the new theory 9P which generates x1',x2',t'
You're then making a conditional error distribution when comparing two rival theories. The 9P theory will minimize this error.
But in no sense can the 9P theory be derived from the initial associative statistical distribution. You are, based on theory (, science, knowledge, etc.) choosing to add a planet, vs. eg., correcting for measurement error, modifiying newton's laws, changing the angle of the earth wrt the solar system... or one of an infinite number of theories which all produce the same error minimization
The sense of "prediction" that science uses (via Popper et al.) is deriving the existence of novel phenomena that do not follow from prior observable distributions.
You want a statistical model that produces a theory of gravity.
Predictions of novel objects derivative from scientific theories arent quantitative data points.
By the way, if you want to see how well gzip actually models language, take any gzipped file, flip a few bits, and unzip it. If it gives you a checksum error, ignore that. You might have to unzip in a streaming way so that it can't tell the checksum is wrong until it's already printed the wrong data that you want to see.
$ curl https://www.gutenberg.org/cache/epub/11/pg11.txt | bzip2 --best | wc
246 1183 48925
Also, ZSTD goes all the way up to `--ultra -22` plus `--long=31` (4GB window— Irrelevant here since the file fits in the default 8MB anyway).