As far as different groups leapfrogging each other for supremacy in various benchmarks, there might be a bit of a "4 minute mile" effect here too - once you know that something is possible then you can focus on replicating/exceeding it without having to worry are you hitting up against some hard limit.
I think the transformer still doesn't get the credit due for enabling this LLM-as-AI revolution. We've had the compute and data for a while, but this breakthough - shared via a public paper - was what has enabled it and made it essentially a level playing field for anyone with the few $B etc the approach requires.
I've never seen any claim by any of the transformer paper ("attention is all you need") authors that they understood/anticipated the true power of this model they created (esp. when applied at scale), which as the title suggests was basically regarded an incremental advance over other seq2seq approaches of the time. It seems like one of history's great accidental discoveries. I believe there is something very specific about the key-value matching "attention" mechanism of the transformer (perhaps roughly equivalent to some similar process used in our cortex?) that gives it it's power.