As far as different groups leapfrogging each other for supremacy in various benchmarks, there might be a bit of a "4 minute mile" effect here too - once you know that something is possible then you can focus on replicating/exceeding it without having to worry are you hitting up against some hard limit.
I think the transformer still doesn't get the credit due for enabling this LLM-as-AI revolution. We've had the compute and data for a while, but this breakthough - shared via a public paper - was what has enabled it and made it essentially a level playing field for anyone with the few $B etc the approach requires.
I've never seen any claim by any of the transformer paper ("attention is all you need") authors that they understood/anticipated the true power of this model they created (esp. when applied at scale), which as the title suggests was basically regarded an incremental advance over other seq2seq approaches of the time. It seems like one of history's great accidental discoveries. I believe there is something very specific about the key-value matching "attention" mechanism of the transformer (perhaps roughly equivalent to some similar process used in our cortex?) that gives it it's power.
It's really not the model, it's the data and scaling. Otherwise the success of different architectures like Mamba would be hard to justify. Conversely, humans getting training on the same topics achieve very similar results, even though brains are very different at low level, not even the same number of neurons, not to mention different wiring.
The merit for our current wave is 99% on the training data, its quality and size are the true AI heroes. And it took humanity our whole existence to build up to this training set, it cost "a lot" to explore and discover the concepts we put inside it. A single human, group or even a whole generation of humans would not be able to rediscover it from scratch in a lifetime. Our cultural data is smarter than us individually, it is as smart as humanity as a whole.
One consequence of this insight is that we are probably on an AI plateau. We have used up most organic text. The next step is AI generating its own experiences in the world, but it's going to be a slow grind in many fields where environment feedback is not easy to obtain.
My take is that prediction, however you do it, is the essence of intelligence. In fact, I'd define intelligence as the degree of ability to correctly predict future outcomes based on prior experience.
The ultimate intelligent architecture, for now, is our own cortex, which can be architecturally analyzed as a prediction machine - utilizing masses of perceptual feedback to correct/update predictions of how the perceptual scene, and results of our own actions, will evolve.
With prediction as the basis of intelligence, any model capable of predicting - to varying degrees of success - will be perceived to have a commensurate degree of intelligence. Transformer-based LLMs of course aren't the only possible way to predict, but they do seem significantly better at it than competing approaches such as Mamba or the RNN (LSTM etc) seq2seq approaches that were the direct precursor to the transformer.
I think the reason the transformer architecture is so much better than the alternatives, even if there are alternatives, is down to this specific way it does it - able to create these attention "keys" to query the context, and the ways that multiple attention heads learn to coordinate such as "induction heads" copying data from the context to achieve in-context learning.
The training set is magical. It took humanity a long time to discover all the nifty ideas we have in it. It's the result of many generations of humans working together, using language to share their experience. Intelligence is a social process, even though we like to think about keys and queries, or synapses and neurotransmitters, in fact it is the work of many people that made it possible.
And language is that central medium between all of us, an evolutionary system of ideas, evolving at a much faster rate than biology. Now AI have become language replicators like humans, a new era in the history of language has begun. The same language trains humans and LLMs to achieve similar sets of abilities.
Are there any Mamba benchmarks that show it matching transformer (GPT, say) benchmark performance for similiar size models and training sets?
Compare to other traditional tech companies… think Uber/AirBnB/Databricks/etc. Their product isn’t an algorithm that a competitor can spin up in 6 months. These companies create real moats, for better or worse, which significantly reduce the ability for competitors to enter, even with tranches of cash.
In contrast, essentially every product we’ve seen in the AI space is very replicable, and any differentiation is largely marginal, under the hood, and the details of which are obscured from customers.
I think we'll see that data, knowledge and intelligence compound and at some point it will be as hard to penetrate as Meta's network effects.
Never say never though - look at Tesla coming out of nowhere to push the big three automakers around! Eventually the established players become too complacent and set in their ways, creating an opening for a smaller more nimble competitor with a better idea.
I don't think LLMs are the ultimate form of AI/AGI though. Eventually we'll figure out a better brain-inspired approach that learns continually from it's own experimentation and experience. Perhaps this change of approach will be when some much smaller competitor (someone like John Carmack, perhaps) rapidly come from nowhere and catch the big three flat footed as they tend to their ginormous LLM training sets, infrastructure and entrenched products.