https://mattmahoney.net/dc/rationale.html
...which leads to the theory of mind known as "compressionism":
https://mattmahoney.net/dc/rationale.html
...which leads to the theory of mind known as "compressionism":
Intelligence _may arise_ from information. Information _necessarily arises_ from compression.
It is very similar, to me, to the concept of humors (https://en.wikipedia.org/wiki/Humorism). Empirically, it was observed that these things had some correlation, and they occurred together, but they are not directly causative. Similarly with the theory of spontaneous generation (https://en.wikipedia.org/wiki/Spontaneous_generation), which was another theory spawned (heh, as it were), from casual causal correlation (again, as it were).
This can be shown rather easily when you show that the time-dynamics of an exemplar system with high information diverge from the time dynamics of a system with high intelligence, i.e. there is a structural component to the information embedded in the system such that the temporal aspects of applied information are somehow necessary for intelligence.
I hope to release some work at least tangentially related to this within the next few years (though of course it is a bit more high-flying and depends on some other in-flight work). If we are to attempt to move towards AGI, I really think we need to stick with the mathematical basics. Neural networks, although they've been made out to be really complicated, are actually quite simple in terms of the concepts powering them, I believe. That's something I'm currently working on putting together and trying to distill to its essence. That will also be at least 1-2 years out, and likely before any temporally-related work, all going well. In my experience, this all is just a personal belief from a great deal of time and thought with the minutia of neural networks. That said, I could perhaps be biased with some notion of simplicity, since I've worked with them long enough for some concepts to feel comfortably second-nature.
To compress data, you have to derive some interesting insight about that data that allows the data to be projected into a lower-dimensional, lower-entropy embedding. (In pop-neurological terms, you have to "make a neuron" that represents a concept, so that you can re-phrase your dataset into a star-schema expressed in its connections to that concept.)
The more different kinds of things a single compressor can do that for — without the compressor itself being made larger — the more "intelligent" the compressor is, no?
Hence, the set vs superset argument.
Kolmogorov actually
Kolmogorov complexity is based directly off of Shannon's information theory. It's the definition of the shortest possible length of a program that can output a particular string.
It is an extension of information theory. If you just take either of the words 'Kolmogorov' or 'complexity' out of the phrase 'Kolmogorov complexity', it loses all meaning. We measure the programs for a given string against that string's Kolmogorov complexity by their entropy. Hence, why at the core, the very pith of the matter at hand is information, which is a measurable quantity. Additionally, K complexity is a lower bound, not a quantity.
This is a similar kind of swap to saying "No, those baked goods are not food, because goods == items bought and sold for currency", if this recontextualizes it a bit.
Zstd, while it might be better than gzip or bzip, is still a very poor compressor compared to an ideal compressor (which hasn't yet been discovered).
That is why zstd acts like a rather bad AI. Note that if you wanted to use zstd as an AI, you would patch out of the source code checksum checks, and you would then feed it a file to decompress (The cat sat on the mat), followed by a few bytes of random noise.
A great compressor would output: The cat sat on the mat. It was comfortable, so he then lay down to sleep.
A medium compressor would output: The cat sat on the mat. bat cat cat mat sat bat.
A terrible compressor would output: The cat sat on the mat. D7s"/r %we
See how each is using knowledge at different levels to generate a completion. Notice also how that few bytes generates different amounts of output depending on the compressors level of world understanding, and therefore compression ratio.
LLMs are sort of unable to do this because they use a fixed tokenizer instead of raw bytes. That means they won't output binary garbage even early on + saves a lot of memory, but it may hurt learning things like capitalization, rhyming, etc we think are obvious.
If I conclude from this that you mean that gzip is intelligent, will I be accused of bad faith?
A highly advanced AI could compress the text and predict the next sequences easily.
This seems like a direct connection like electricity and magnetism.
And maybe that's why English needs to be about 1 bit because we're not very intelligent.