I ask as a noob in this area.
I ask as a noob in this area.
- Unsupervised learning techniques, e.g. transformers and diffusion models. You need unsupervised techniques in order to utilize enough data. There have been other unsupervised techniques in the past, e.g. GANs, but they don't work as well.
- Massive amounts of training data.
- The belief that training these models will produce something valuable. It costs between hundreds of thousands to millions of dollars to train these models. The people doing the training need to believe they're going to get something interesting out at the end. More and more people and teams are starting to see training a large model as something worth pursuing.
- Better GPUs, which enables training larger models.
- Honestly the fall of crypto probably also contributed, because miners were eating a lot of GPU time.
You're right that conditional generation start to blur the lines though.
While you're right about GANs, diffusion models as transformers as transformers are most commonly trained with supervised learning.
Unsupervised is a confusing term as there is always an underlying loss being optimized and working as a supervision signal, even for good old kmeans. But generative models are generally considered to be part of unsupervised methods.
Exactly. The growth in the next decade is going to be unimaginable because now governments and MNCs believe that there realistically be progress made in this field.
Of course those two releases didn't fall out of the sky.
There’s been open source AI/ML for 20+ years.
Nothing comes close to the massive milestones over the past year.
Take AlexNet - the major "oh shit" moment in image classification.
It had an absolutely mind-blowing number of parameters at a whopping 62 million.
Holy shit, what a large network, right?
Absolutely unprecedented.
Now, for language models, anything under 1B parameters is a toy that barely works.
Stable diffusion has around 1B or so - or the early models did, I'm sure they're larger now.
A whole lot of smart people had to do a bunch of cool stuff to be able to keep networks working at all at that size.
Many, many times over the years, people have tried to make larger networks, which fail to converge (read: learn to do something useful) in all sorts of crazy ways.
At this size, it's also expensive to train these things from scratch, and takes a shit-ton of data, so research/discovery of new things is slow and difficult.
But, we kind of climbed over a cliff, and now things are absolutely taking off in all the fields around this kind of stuff.
Take a look at XTTSv2 for example, a leading open source text-to-speech model. It uses multiple models in its architecture, but one of them is GPT.
There are a few key models that are still being used in a bunch of different modalities like CLIP, U-Net, GPT, etc. or similar variants. When they were released / made available, people jumped on them and started experimenting.
SDXL is 6.6 billion.
The availability of GPU compute time. Up until the Russian invasion into Ukraine, interest rates were low AF so everyone and their dog thought it would be a cool idea to mine one or another sort of shitcoin. Once rising interest rates killed that business model for good, miners dumped their GPUs on the open market, and an awful lot of cloud computing capacity suddenly went free.
Emad Mostaque and his investment in stable diffusion, and his decision to release it to the world.
I'm sure there are others, but those are the two that stick out to me.
Not that more advances don't happen with sustained hype, just there's some sort of tipping point involving usefulness based either on improvement of the thing in question or it's utility elsewhere.