The data fed into these systems is just measurements of target systems (eg., of light for photographs). This data is radically incomplete, so no compression of it will be an accurate (eg., 3d) model.
To reconstruct the world you need to measure the measuring device (ie., the body) as it interacts with the target system. In most cases you need, also, the hypothesis behind the action to resolve ambiguity in the measurement data.
Eg., you need to know you moved your hand to touch the fireplace to properly interpret what 'finger pain' means.
It is for reasons of this kind that 'comprehension' is an engineering problem, not a programme for a universal turing machine (which has no device boundaries).
It is engineering in the sense that 'compression' has to occur under the right measurement procedures, with hypothesis-laden action, etc.
Focusing on mathematical abstracta completely misses the problem.
All deep systems do is compress their training data to form archetypes and compare novel input to compressed archetypes. Since the data itself is necessarily profoundly ambiguous there is only a trivial sense of 'generalisation' achievd.