There is definitely a use for this kind of model, but this also shows why Transformers are still the main architecture we use today.
Newer methods are going back to similar concepts but trying to get past previous bottlenecks given what we've learned since then about transformers.
Beware of feel-good traps like this.
If you were to map the human connectome to a computational neural network down to the ion channel, it'd be at least 500 quadrillion parameters*. That's at least 5-6 orders of magnitude beyond what is currently possible with SOTA ML which means that even if the human brain was 99% devoted to compressing those tokens, that 1% that could actually do work with them is still a thousand times bigger than GPT4. There be emergent dragons.
* This is a fascile argument to begin with since biological neuron signals aren't quantized and ion channels are far too complex to map to a single static parameter
Perhaps, but they inherently have imprecision due the nature of being biological. Expose the same neuron to the same input N times, and you'll get a range of output. The effect of noisy analog data is, broadly speaking, similar to the effect of low-resolution digital data.
but also, on a broader scale, if a transformer model is presented with a long input that does not fit in its context (e.g.: you are building a chatbot, and you have a very long chat history), it must "compress" or "forget" some of that information (e.g.: repeatedly summarizing historical messages, dropping them and appending the summary at the beginning of the input).
Mamba/RWKV/other "recurrent" architectures, can theoretically operate on unbounded input lengths; they "forget" information from earlier tokens over time, but is that not comparable to what a transformer must do with input lengths greater than their context window?
Your assumption that the same effort could be given to MAMBA or RWKV isn’t necessarily true. I’m sure there’s research arms looking to see if they scale to the level OpenAI or Google have their transformers at now, but they’re still very much in their infancy.
This all ignores the risk of them not panning out and the lost time not focusing some energy on transformers.
New technologies aren't generally just created out of whole cloth by some genius, they're build up of layers of incremental improvements, each of which is modest in its own right but which when taken together are groundbreaking.
[I would be happy to say more but I just got a top-level comment flag killed, presumably because they though I was advertising, so I won't mention the company name.]
Probably because the pace of development is now no longer determined by ML researchers but by CS engineers. ML is "easy to pick up" until you have to understand the reasons certain architectures work or don't for different applications.
OpenAI's moat is the high cost of training as you mentioned but this might be obsolete in a few papers down the road.
They themselves realized this by trying to turn it into a platform but I don't think it's enough.