The keyword here is cheap. Any sufficiently powerful maximizer with an infinite horizon __has__ to develop curiosity otherwise they will not be able to maximise their reward function.
In-fact, I'd argue that this is true for most mammalian functions such as taking care of our pack, exhibiting pro-social behaviour and so on, but there is a caveat. For this to happen, there needs to be an actual benefit in the behaviour.
Due to evolution operating mostly linearly with small changes through genetic and epigenetic information passing, there seems to be relatively little variation between generations, which then implies that it is difficult for some candidates to overwhelm everyone else in a winner takes all fashion, hence maximization of replication will eventually result in cooperation simply because that allows genes in support of it to continue replicating, effectively self-selecting for itself.
We saw this in the OpenAI video where agents eventually learned to cooperate in what was effectively a prisoners dilemma. In the video, there were two teams, the hiders and the seekers, in an environment that could be manipulated. Eventually the two teams learned strategies. From the perspective of seekers, their utility function involved observing the hiders. For the hiders, their utility function was to minimize their exposure to the seekers. For all intents and purposes, and for each team, the other agents were part of the environment. So given two hiders, one could hide behind the other to minimize their exposure to the seekers. This is effectively a defect. Eventually however, the hiders learn to cooperate and instead cooperatively manipulate the environment through strategies.
---
Sorry if I got a little bit off topic there.
Regardless, what causes us to learn is neurotransmitters getting released because certain circuits activate, the neurotransmitters charge a neuron which causes it to fire. Connections between neurons get reinforced if they are frequently used, which reinforces them, makes them cheaper, and that inevitably reinforces certain patterns of behaviour.
---
What i propose is that we should instead look into analogies as a means of learning. Humans seem to be great at using analogies. Mathematically speaking, an analogy is a functor between categories. A category is a collection of objects and morphisms (directed relationships) between the objects, this is as abstract and simple as it gets. A functor between categories essentially maps the objects and the morphisms of one category to the respective of the other. When we use an analogy, we do the same.
> A is to B as X is to Y
This then allows us to learn the morphism in the category with {A,B} using just known relationships.
I think I got off topic again, but this is something that I have been recycling in my head for a while and needed to eventually get out.