I'm kind of shocked. I thought there would be more dynamism by now and I stopped dabbling in like 2018.
I'm kind of shocked. I thought there would be more dynamism by now and I stopped dabbling in like 2018.
This particular (tock) is still playing out. The next (tick) does not feel imminent and will likely depend on when we discover the limits of the transformers when it comes to solving for long tail of use-cases.
My $0.02.
There are certainly tradeoffs to both, the general transformer motif scales very well on a number of axis, so that may be the dominant algorithm for a while to come, though almost certainly it will change and evolve as time goes along (and who knows? something else may come along as well <3 :')))) ).
I think it's dominance is not going to substantially change any time soon. Dont you know, the solution to all leetcode interviews is a hash table?
Furthermore, I think a replacement will require that we _understand_ what the current crop of models are doing mechanically. Some of it was motivated in [1].
[1] https://openaipublic.blob.core.windows.net/neuron-explainer/...
Personally, I think going linear instead of quadratic for a core operation that a system needs to do is by definition an optimization.
My bet will be on something else than gradient descent and backprop but really I don't wish any company or country to reach agi or any sophisticated ai ...
If a 100x improvement in performance is left on the table, then surely even lower priority optimizations won't be implemented any time soon.
Consider this: a lot of clever attention optimizations rely on some initial pass to narrow the important tokens down and discarding them from the KV cache. If this was actually possible, then how come the first few layers of the LLM don't already do this numerically to focus their attention? Here is the shocker: they already do, but since you're passing the full 8k context to the next layer anyway, you're wasting it on mostly... Nothing.
I repeat: Does the 80th layer really need the ability to perform attention over all the previous 8k outputs of the 79th layer? The first layer? Definitely. The last? No. What happens if you only perform attention over 10% of the outputs of layer 79? What speedup does this give you?
Notice how the model has already learned the most optimal attention scheme. You just need to give it less stuff to do and it will get faster automatically.
Step change, then optimization of that step change
Kind of like a grand father clock with a huge pendulum swinging to one side, then another(commonly used metaphor).
There seems to be no rhyme or reason, no scientific insight, no analysis. They just try a million different permutations, and whatever scores the highest on the benchmarks gets published.
E.g. "In-context Learning and Induction Heads" is an excellent paper.
Another paper ("ROME") https://arxiv.org/abs/2202.05262 formulates hypothesis over how these models store information, and provide experimental evidence.
The thing is, a 3-layer MLP is basically an associative memory + a bit of compute. People understand that if you stack enough of them you can compute or memorize pretty much anything.
Attention provides information routing. Again, that is pretty well-understood.
The rest is basically finding an optimal trade-off. These trade-off are based on insights based on experimental data.
So this architecture is not so much accidental as it is general.
Specific representations used by MLPs are poorly understood, but there's definitely a progress on understanding them from first principles by building specialized models.
I wonder shouldn't AI be the best tool to optimize itself?
There's still some room for experimenting if you care about memory/power efficiency, like MoE models, but they're not as well understood yet.
So until the hardware allows for comparable (say with 2-4x) thoroughput of samples per second I expect model architecture to mostly be static for most effective models and dynamic architectures to be an interesting side area.