Alas the goal isnt the same, and here "predict" is the culprit. What ML systems do is offer statistical estimates, which is one very narrow meaning of "predict".
When I "predict" what will happen when a glass falls, I'm not giving you any statistical estimates. I'm running a mental simulation of the world in its relevant parts and using this to predict what will happen.
The heart of basically almost all knowledge of the world is actually counterfactual (ie., requiring simulation), it's always "if this happens (in the simulation), this will happen (in the simulation)".
To acquire this ability to predict one has to understand, ie., to have the right concepts, reason with them and so on. When a person "has the concept `7`" they can reason thus: "6 eggs would be fewer", "7 is an odd number", "7 is a whole quantity", "with 7 less of 10, i'd have 3" etc.
Ie., `7` is available to play a fundamental role in simulating scenarios and reasoning with them.
When a NN produces a mapping from {Images} -> {Labels} such that the squiggle `7` produces the label `"7"` it doesnt do so via prediction in the meaningful sense of the world. Rather it has found a means of averaging sqiggles, and compares one squiggle to its averages and reports what the historical similar squiggles were labelled.
This is a game which is only useful because we, who can reason with `7` profoundly about the world (etc.), have produced datasets which express our knowldge of the world in ways that dont require it. Ie., we present to the machine something whose average is useful.
The machine is not trying to, nor does it, nor can it, predict. It has no knoweldge at all. It is just a correspondance table we have provided.