But that's not a question of fairness, rather it's a question of feasibility. If it took us many thousands of years to learn our background knowledge from the real world over many human generations, it's difficult to see how we can reproduce this result with the comparatively poor computational resources and data in our disposal.
There is a peculiar double-blindness in machine learning today, I think, where people are hoping to learn extremely difficult concepts, like meaning in language or like all of intelligence, from simultaneously too much and too little data. Too much because humans don't need to train on the entire web to learn meaning (and Large Language Models trained on the entire web still don't learn it). And too little because if you think of the complexity of the real world and the amount of information that we take in with our senses just sitting still looking around, this is an amount of information that can simply not be matched by the largest imaginable dataset that we could create.
So what's the altnerative? I have a parable (oh no). What do you do when you need a fire? Well, clearly, you light a fire, maybe with matches or with a lighter etc. That's because you know how to light a fire and because the implements to do so are now cheap commodities that most humans can afford easily (I bet even Kalahari bushmen use BIC ligthers nowadays...). What you certainly don't do is sit around waiting for a fire to occur naturaly, say by thunder strike, like humans presumably did before discovering how to make fire from scratch. Because that could take ages and because you have the knowledge necessary to not have to wait for ages. In the same way, we could wait around for ages trying to train systems to develop complex abilities like understanding or intelligence from ever lager datasets- which can take many decades, since, like I say, we have simultaneously too little and too much data; or, we can find a way to transfer the background knowledge bestowed upon us by thousand years of evolution to guide the training of our learning systems towards the goals we want them to achieve, whatever those are. We can give them the spark to start a fire. Or maybe we can't. But, if we can, then there is no sensible reason why we shouldn't.
To clarify, I'm not saying we should go back to feature engineering. Feature engineering was necessary in the past because there is no good way to imbue neural networks with background knowledge. Notably, it's not possible to use a trained neural net as a feature of another neural net, so it's not possible to build up from low-level concepts to higher-level concepts unless it's done end-to-end in the same model, which is limiting, and yet another reason for the gigantism of neural net training datasets. I don't know what the solution is, though. Clearly not explicitly coding expert knowledge in production rules as in expert systems. Much of our knowledge is maybe impossible to articulate explicitly. So we must find a way to encode implicit knowledge, also.
But, again, we don't have to encode _everything_. We can find good predicates, in Vapnik's terminology, and then let the learning systems do the rest. But that can only work _if_ our learning systems _can_ do the rest.
Yes, translational invariance in CNNs is a good example. But it's still not the whole story.
>> Is choosing suitable starting weights/architecture/"predicates" by hand-designing based on our own built up information qualitatively any different? It still seems like "hiding" utilization of a huge amount of background knowledge about digits/symbols/images/reality.
It's basically a trade-off. If you have good background knowledge, you don't need a lot of data. Good background knowledge helps you build robustly generalisable concepts. And if you can reuse the learned concepts as background knowledge, then the sky is the limit. But, if you don't have background knowledge, you need to make up for it, and the only way we know is to train on lots and lots of data- with the limitations that involves (overfitting, large computational costs, etc).
P.S. Sorry- this comment is a bit sloppy and hence overlong.