but also, on a broader scale, if a transformer model is presented with a long input that does not fit in its context (e.g.: you are building a chatbot, and you have a very long chat history), it must "compress" or "forget" some of that information (e.g.: repeatedly summarizing historical messages, dropping them and appending the summary at the beginning of the input).
Mamba/RWKV/other "recurrent" architectures, can theoretically operate on unbounded input lengths; they "forget" information from earlier tokens over time, but is that not comparable to what a transformer must do with input lengths greater than their context window?
There is definitely a use for this kind of model, but this also shows why Transformers are still the main architecture we use today.
Newer methods are going back to similar concepts but trying to get past previous bottlenecks given what we've learned since then about transformers.
Beware of feel-good traps like this.
If you were to map the human connectome to a computational neural network down to the ion channel, it'd be at least 500 quadrillion parameters*. That's at least 5-6 orders of magnitude beyond what is currently possible with SOTA ML which means that even if the human brain was 99% devoted to compressing those tokens, that 1% that could actually do work with them is still a thousand times bigger than GPT4. There be emergent dragons.
* This is a fascile argument to begin with since biological neuron signals aren't quantized and ion channels are far too complex to map to a single static parameter
Perhaps, but they inherently have imprecision due the nature of being biological. Expose the same neuron to the same input N times, and you'll get a range of output. The effect of noisy analog data is, broadly speaking, similar to the effect of low-resolution digital data.