I am myself trying to get closer to a clear formulation of this problem, which is why I'm writing here. Here's what I have so far,
ML systems (eg., NN) remember averages (, compressions) of historical data. They are useful, wrt the problem, iif (1) the problem's target function exists; (2) the data is relevant, unambiguous, well-carved; and if (3) these properties will hold regardless of likely permutations to the problem's framing.
Systems are given data with these properties by significant amounts of experimental design, work, and effort by people. Absent these properties, data is useless.
Producing data with these properties requires intelligence, and no machine systems exist which can do it.
My issue with research into ML
on the whole, is that it *assumes* these properties
and then explains how the systems work. I understand why this is interesting from a formal perspective... but it fails to note that this situation is almost never how ML is used.
There is no function from Image->Animal, ie., biologists arent just idiots who could have just looked at some pixel patterns. Pixel patterns are radically ambigious wrt to `Animal`, and so even an infinite sampling of (Image, Animal) is not enough for ML.
... so what on earth are ML systems doing?
This is a bigger research question: to characterise how ML performs when this assumed setup fails. And you know, that research almost doesnt exist. This is an industry led by partisans to its success.
What do you think would happen if research actually talked about the dynamics of ML systems performance when (1) the target doesnt exist; (2) data isnt relevant & unambigious; (3) the problem framing will permute most times its deployed....
Suddenly we'd have an explaination of why 2016 wasnt the year self-driving cars were delivered. And indeed, likewise, of why even 2036 wont be.