I am beginning this journey, and it takes a lot of the magic out of it.
Really what is happening is just mathematical operations: a lot of matrix multiplications, some simple non-linear functions (because linear algebra alone cannot represent all logic), some statistical stuff to stop the numbers getting out of control.
Importantly there are millions or billions of "magic numbers" that get updated as the model learns called parameters.
Whether you train the network on representation A of the image, or representation B of the image, doesn't change much philosophically. It is just a function f(x) applied before the function g(x, parameters), so you get g(f(X), parameters).
Now you could argue, well it might mean "something". After all you can reduce our brains to being a big mathematical functions.
Possibly.
But I think in this case it is too simple. The AI is just looking at a slightly different representation. Similar to changing VS code from/to dark mode for humans, or something like that. And the modellers might be changing representations all the time anyway as they tinker with things.