Remixing Jazz tokens might never generate classical music, but an evolutionary algorithm using musical notes and basing its cost function on what humans seem to enjoy might still discover something similar or better. A transformer could then generate infinite versions of that discovery.
More interesting is that these tribes also tend to lack the ability to recall even basic quantities of things if it's more than 2. You can give them 5 balls, but they will have extremely poor recollection of the number after just a few moments. There's these countless things humans need to simply invent, from nothing, over and over to advance to the next technological era. It only seems like a natural application of logic in hindsight, because it actually works!
[1] - https://www.sciencedaily.com/releases/2012/02/120221104037.h...
Of corse one in a million that some will fall into a different evolutionary path.
The number on the success side is statistically way too significant so you can't revolve anything with this one.
I am not saying it's not superior (I do believe it is), just that this does not really support it.
Basically, a single human acting on its own during their single lifetime has almost zero chance of inventing mathematics (as evidenced by the tribed you mention, but also examples of children lost in woods and raised by animals).
And yes, it's going to be a slow process. Human intelligence is not about speed. The capability to do simple arithmetic rapidly correlates with intelligence in humans, yet a calculator can perform said calculations many millions of times faster than the fastest human. Yet that of course doesn't mean said calculator is thus millions of times more intelligent. The question is how long would it take an LLM to discover mathematics starting from the same empty baseline? And, using current methods, the answer is quite simple: never.
https://earthsky.org/earth/fish-can-do-math-cichlids-stingra...
To see how silly these things are imagine somebody claiming to have taught mathematics to one of these numberless tribes by training them to pick a larger group of items when it's blue, and pick a smaller group of items when it's yellow. Somebody claiming that is teaching them math would obviously be looked at like an idiot. It's a circus trick - and a rather apt reflection of what passes for research now a days.
Infinite recursion has to be implemented by iteration in the physical universe.
And then just feed it endless continuous video of the world from a persons perspective (ie, not just random jumps around scenes) How is that any different to the way humans developed our own theories and understanding on the universe we found ourselves in? Obviously a video feed isn't quite the same as having a myriad of senses like we have for interfacing with the world. But theres no reason to say newtons laws need touch or smell to be figured out (maybe that is a bad example for not requiring touch, being about mass and force, but you get the point). Learning language and understanding of the world from scratch might depend on having physcial interaction with the world, sure, but thats again just different sets of input training data that a multimodal model can make connections between if there were pressure to do so.
Then you simply prompt the LLM to define the laws it has observed in a way that can use language to transmit the idea. In the human equivalent, a "prompt" doesn't need to be (but can be) a direct question being asked. It could just be our urge to share ideas, an emotion perhaps that moves us to action. I don't want to go too far on this thought as it is pure speculation without serious scientific backing, but for the sake of the thought experiment with all this in mind you could even say that brains are much like an LLM that is in a constant state of training & updating, but also being "prompted" from our environment and biological feedback loops being fed in. I don't feel like this idea is even that controversial or remotely original.
The difference is that models of today are already trained on the corpus of human knowledge so they already know all of this instead of having to figure it out. But I don't really see an argument for why an LLM wouldn't be able to crystallise observed patterns into mathematical formulae if it had the appropriate influences that would encourage it to do so. Its just that we don't care to train them inefficiently. I imagine this will be tried in the near future though.