> If the argument to get a new keyboard is: "i like it", then this should suffice
This seems like exactly what LLMs are supposed to be good at, according to you, so why don't they just near-losslessly compress the data first, and then train on that?
Also, if they're so good at this, then why are their answers often long-winded and require so much skimming to get what I want?
I'm skeptical LLMs are accurately described as "near lossless de/compression engines".
If you change the temperature settings, they can get quite creative.
They are their algorithm, run on their inputs, which can be roughly described as a form of compression, but it's unlike the main forms of compression we think of - and it at least appears to have emergent decompression properties we aren't used to.
If you up the lossy-ness on a JPEG, you don't really end up with creative outputs. Maybe you do by coincidence, and maybe you only do with LLMs - but at much higher rates.
Whatever is happening does not seem to be what I think people typically associate with simple de/compression.
Theoretically, you can train an LLM on all of Physics, except a few things, and it could discover the missing pieces through reasoning.
Yeah, maybe a JPEG could, too, but the odds of that seem astronomically lower.