People figuring out how to train and scale newer architectures (like transfomers) effectively, to be wildly larger than ever before.
Take AlexNet - the major "oh shit" moment in image classification.
It had an absolutely mind-blowing number of parameters at a whopping 62 million.
Holy shit, what a large network, right?
Absolutely unprecedented.
Now, for language models, anything under 1B parameters is a toy that barely works.
Stable diffusion has around 1B or so - or the early models did, I'm sure they're larger now.
A whole lot of smart people had to do a bunch of cool stuff to be able to keep networks working at all at that size.
Many, many times over the years, people have tried to make larger networks, which fail to converge (read: learn to do something useful) in all sorts of crazy ways.
At this size, it's also expensive to train these things from scratch, and takes a shit-ton of data, so research/discovery of new things is slow and difficult.
But, we kind of climbed over a cliff, and now things are absolutely taking off in all the fields around this kind of stuff.
Take a look at XTTSv2 for example, a leading open source text-to-speech model. It uses multiple models in its architecture, but one of them is GPT.
There are a few key models that are still being used in a bunch of different modalities like CLIP, U-Net, GPT, etc. or similar variants. When they were released / made available, people jumped on them and started experimenting.