Actually, algorithmic information theory shows that maximum compression necessarily entails maximum comprehension, because the only way to maximally compress something is to exactly understand the process producing it.
Actually, algorithmic information theory shows that maximum compression necessarily entails maximum comprehension, because the only way to maximally compress something is to exactly understand the process producing it.
They're making a practical distinction that you generally don't have access to the actual thing in an empirical format for which compression will achieve true learning. Instead you have access to training data which represents, let's say, a projection of the actual thing in a smaller space with fewer dimensions.
Like trying to learn from images instead of the 3d world. Humans learn to distinguish between objects in a 3-dimensional space using sight and interaction. This learning generalizably transfers to recognition in 2 dimensions. We don't generally equip models with robotic interfaces to train in 3d before benchmarking them on ImageNet.
Don't they train models using 3D rendering and simulations ? We have relatively realistic simulations for various scenarios - having a learned model that could make inferences based on those complex simulations sounds like a win.
If we use human "comprehension" as a reference point, then the relevant point of comparison should be the understanding a human can develop given the same inputs.
Most ML problems are things humans are quite good at and have a lot of context to draw from.
A string X maximally compresses datsets Y iff X is a 'comprehending' of Y.
'OK'... but what produces and evaluates X? ie., comprehension.
This is the problem with defining these terms mathematically; you state the problem in basically useless ways.
Yes, you can specify what eqn produces the mass of the higgs boson. Thats basically no guide to building the LHC.
The production of such understanding is not abstract. Comprehension isnt a relation between two binary strings; it is an action taken in an environment with a goal.
It uses Kolmogorov complexity, which is defined as the length of a shortest computer program in a predetermined programming language that produces the object as output. Note that this measure is relative to a programming language, not a program, and the exact choice of language doesn't matter too much. Compression means creating a smaller program that produces the same output, and to produce the same output with less code necessarily requires more understanding.
As a concrete example, imagine the output is [1,2,fizz,4,buzz,fizz,7,8,fizz,buzz,11,fizz,13,14,fizzbuzz..1000). The longest program to output this would just hard-code it in the source code (much as a very inexperienced programmer might solve the problem, or a large neural net). Someone with a better understanding would write a program using iteration and the modulus operator, which would be shorter.
One could argue it's not about compression in bits but compression to primitives that makes sense to the human mind. But then the definition becomes to fuzzy because it naturally invites the question "Who's mind?"
It’s unclear why compression is necessary there, except as a practical benefit.