I’m sure the answer is in there somewhere but I’m not well versed in AI and I simply can’t figure it out.
I’m sure the answer is in there somewhere but I’m not well versed in AI and I simply can’t figure it out.
The model's job is to predict what comes next. The entropy coder's job is to encode the difference between the prediction and what actually comes next so that the most likely outcome uses as few bits as possible. The more accurate the model is, the less the difference between reality and prediction, the less bits the entropy coder needs and the better the compression.
Simple compression algorithms have simple models, like "if I see the same byte 10 times, the 11th is likely to be the same". But you can also use a LLM as your model, as completing text with the most likely word is what LLMs do.
Here they did the opposite. Instead of using a model for compression, by using a few tricks, they used a compression algorithm as a model: the most likely outcome is when the compression algorithm uses less bits to encode the result. And the original authors have shown that, in some tasks, the simple model that can be extracted out of gzip beats much more complex LLMs.
Embeddings use fixed-size vectors to minimize the dot product between vectors of similar inputs. Compressors use a variable-length encoding to minimize the overall stored size.
If any byte sequence is a correct file (unlikely, but mostly because compression algorithms try to be robust against corruption), then this is easy to reverse, you just generate a random sequence of bytes and then decompress it.
Basically you can turn a compression algorithm into a probability distribution by inserting random bytes wherever the decompression algorithm tries to read one, but sometimes not all bytes are allowed.
You can then reason about this probability distribution and see what it's properties are. Typically something with a probability of 'p' will require -log(p)/log(2) bits.
For compression, word sequences that have higher probability should be encoded with shorter codes, so there is a direct relationship. A well known method to construct such codes based on probabilities is Huffman coding.
This works whether you use a statistical language model using word frequencies or an LLM to estimate probabilities. The better your language model (lower perplexity) the shorter the compressed output will be.
Conversely, you can probably argue that a compression algorithm implicitly defines a language model by the code lengths, e.g., it assumes duplicate strings are more likely than random noise.
If you compress `ABC`, it will be X bytes. If you then compress `ABCABC`, it will not take 2x bytes. The more similar the two strings that you concatenate, the less bytes it will take. `ABCABD` will take more than `ABCABC`, but less than `ABCXYZ`.
BERT is, by todays standards, a very small LLM, which we know has weaker performance than the billion-param scale models most of us are interacting with today.
Heh. So does that make it a MLM (medium)?
I've always found it funny that we've settled on a term for a class of models that has a size claim... Especially given how fast things are evolving...
Highly recommend reading Ted Chiang's "ChatGPT Is a Blurry JPEG of the Web"[0] to get a better sense of this.
Keeping this fact in your mental model neural networks can also go a long way to demystify them.
0. https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
Memories are always part fabrications. You can't return to previous mental states (you only think you do) and we have no real clue what really informs decisions i.e preferences shape choices just as much as choices shape preferences.
Your brain will happily fabricate rationals you sincerely believe for decision that couldn't possibly be true i.e split brain experiments
Plain english TL;DR:
- if you limit your task to binning snippets of text
- and the snippets are very well-defined (ex. code vs. Filipino text)
- the snippets _bytes_ could be compared and score well, no text understanding needed
- size delta of a GZIP after adding one more sentence acts as an ersatz way to compare sets of bytes to eachother (ex. you can imagine a GZIP containing 0xFFABCDEF that has 0xFFABCDEF added to it will have a size delta of 0)
There's two issues:
- I don't have an ear for what's simple vocabulary versus tryhard, I go into a mad loop when I try
- even if I actively notice it, substitution can seem very far away from intent. Simple wouldn't have occurred to me - I wanted to say something more akin to sloppy / stunt and ersatz is much closer to "hacky" in meaning than simple. Think MacGyver.
But I should do the exercise of at least scanning for words more often and aim for wider audience - I would have known ersatz was an outlier and I shouldn't feel it's condescending or diluting meaning, it's broadening the audience who can parse it