Actually we have an awful lot of those.
I'm not sure if emergent is quite the right term here. We carefully craft a scenario to produce a usable gradient for a black box optimizer. We fully expect nontrivial predictions of future state to result in increasingly rich world models out of necessity.
It gets back to the age old observation about any sufficiently accurate model being of equal complexity as the system it models. "Predict the next word" is but a single example of the general principle at play.
This is admission we don't know how it emerges.
Sure, we expect the behavior to emerge, but we don't know how.
The "black box" bit refers to a generic, interchangeable optimization algorithm that simply makes the number go down (or up or whatever).
There are certainly various details about the internal workings of models that we don't properly understand but a blanket claim about the whole is erroneous.
[1] https://en.wikipedia.org/wiki/What_Is_It_Like_to_Be_a_Bat%3F
We fully do. There is a significant quality difference between English language output and other languages which lends a huge hint as to what is actually happening behind the scenes.
> but how exactly does anthill behavior come from ant behavior?
You can't smell what ants can. If you did I'm sure it would be evident.
1. Can you reveal "what's actually happening behind the scenes" beyond the hint you gave? I can't figure it out.
2. Can you explain how an ants sense of smell leads to anthills?
Ant 0: doesn’t seem to be dangerous here. I’ll drop a scent.
Ant 1: oh cool, a safe place. And I didn’t die either. I’ll reinforce that.
Ant 142,857,098,277: cool anthill.
?
Not your fault obviously, but they have not yet described what that huge hint is, and I'm just at the edge of my seat with anticipation here.