Yes, humans come into the world with seemingly a very large amount of background
knowledge that we can then use to learn new concepts from very few examples. And
as you say this is probably the result of many thousands of years of evolution.
But that's not a question of fairness, rather it's a question of feasibility. If
it took us many thousands of years to learn our background knowledge from the
real world over many human generations, it's difficult to see how we can
reproduce this result with the comparatively poor computational resources and
data in our disposal.
There is a peculiar double-blindness in machine learning today, I think, where
people are hoping to learn extremely difficult concepts, like meaning in
language or like all of intelligence, from simultaneously too much and too
little data. Too much because humans don't need to train on the entire web to
learn meaning (and Large Language Models trained on the entire web still don't
learn it). And too little because if you think of the complexity of the real
world and the amount of information that we take in with our senses just sitting
still looking around, this is an amount of information that can simply not be
matched by the largest imaginable dataset that we could create.
So what's the altnerative? I have a parable (oh no). What do you do when you
need a fire? Well, clearly, you light a fire, maybe with matches or with a
lighter etc. That's because you know how to light a fire and because the
implements to do so are now cheap commodities that most humans can afford easily
(I bet even Kalahari bushmen use BIC ligthers nowadays...). What you certainly
don't do is sit around waiting for a fire to occur naturaly, say by thunder
strike, like humans presumably did before discovering how to make fire from
scratch. Because that could take ages and because you have the knowledge
necessary to not have to wait for ages. In the same way, we could wait around
for ages trying to train systems to develop complex abilities like understanding
or intelligence from ever lager datasets- which can take many decades, since,
like I say, we have simultaneously too little and too much data; or, we can find
a way to transfer the background knowledge bestowed upon us by thousand years of
evolution to guide the training of our learning systems towards the goals we
want them to achieve, whatever those are. We can give them the spark to start a
fire. Or maybe we can't. But, if we can, then there is no sensible reason why we
shouldn't.
To clarify, I'm not saying we should go back to feature engineering. Feature
engineering was necessary in the past because there is no good way to imbue
neural networks with background knowledge. Notably, it's not possible to use a
trained neural net as a feature of another neural net, so it's not possible to
build up from low-level concepts to higher-level concepts unless it's done
end-to-end in the same model, which is limiting, and yet another reason for the
gigantism of neural net training datasets. I don't know what the solution is,
though. Clearly not explicitly coding expert knowledge in production rules
as in expert systems. Much of our knowledge is maybe impossible to articulate
explicitly. So we must find a way to encode implicit knowledge, also.
But, again, we don't have to encode _everything_. We can find good predicates,
in Vapnik's terminology, and then let the learning systems do the rest. But that
can only work _if_ our learning systems _can_ do the rest.
Yes, translational invariance in CNNs is a good example. But it's still not the
whole story.
>> Is choosing suitable starting weights/architecture/"predicates" by
hand-designing based on our own built up information qualitatively any
different? It still seems like "hiding" utilization of a huge amount of
background knowledge about digits/symbols/images/reality.
It's basically a trade-off. If you have good background knowledge, you don't
need a lot of data. Good background knowledge helps you build robustly
generalisable concepts. And if you can reuse the learned concepts as background
knowledge, then the sky is the limit. But, if you don't have background
knowledge, you need to make up for it, and the only way we know is to train on
lots and lots of data- with the limitations that involves (overfitting, large
computational costs, etc).
P.S. Sorry- this comment is a bit sloppy and hence overlong.