What are the resources needed to train this model?
If someone just gave you the model for free, what resources would you need to use it to generate new results?
What are the resources needed to train this model?
If someone just gave you the model for free, what resources would you need to use it to generate new results?
This means the parameters of the trained model fit in something like 7GB (decoder only, half-precision floats) to 24GB (full model, full-precision). To actually run the model, you will need to store those parameters, as well as the activations for each parameter on each image you are running, in (video) memory. To run the full model on device at inference time (rather than r/w to host between each stage of the model) you would probably want an enterprise cloud/data-center GPU like an NVIDIA A100, especially if running batches of more than one image.
The training set size is ~97TB of imagery. I don't think they've shared exactly how long the model trained for, but the original CLIP dataset announcement used some benchmark GPU training tasks that were 16 GPU-days each. If I were to WAG the training time for their commercial DALL-E 2 model, it'd probably be a couple of weeks of training distributed across a couple hundred GPUs. For better insight into what it takes to train (the different stages/components of) a comparable model, you can look through an open-source effort to replicate DALL-E 2.[2]
[0] https://cdn.openai.com/papers/dall-e-2.pdf [1] https://openai.com/blog/clip/ [2] https://github.com/lucidrains/dalle2-pytorch
> you would probably want an enterprise cloud/data-center GPU like an NVIDIA A100, especially if running batches of more than one image.
That doesn't seem so bad.
looks up price of NVIDIA A100 - $20,000
oh...ok I'll probably just pay for the service then
For the half-precision version at 7GB there are a ton more options (the RTX 3060 has 12GB for example at ~$450).
I do hope that the conversation starts to acknowledge the difference between sunk costs and running costs.
Employees, office leases and equiment are all happening, regardless and ongoing.
Training DALL-E 2: very expensive, but done now. A sunk cost where every dollar coming in makes the whole endeavor more profitable.
Operating the trained model: still expensive, but you can chart out exactly how expensive by factoring in hardware and electricity.
I believe that by not explicitly separating these different columns when discussing expense vs profit, we're making it harder than it needs to be to reason about what it actually costs every time someone clicks Generate.
But I seem to remember they were running 1,000+ 32gb GPUs for 3 months to train it and keeping that infrastructure running day-to-day and tweaking parameters as training continued was the bulk of the 100 pages. It is beyond the reach of anybody but a really big company, at least in the area of very large models, and the large models are where all the recent results are. I wish I was more bullish on algorithm improvements meaning you can get better results on less hardware; there will definitely be some algorithm improvements, but I think we might really need more powerful hardware too. Or pooled resources. Something. These models are huge.
Is https://github.com/facebookresearch/metaseq/blob/main/projec... what you're referring to?
Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months.
Based on this blog post where they scale to 7,500 'nodes', they say:
> A large machine learning job spans many nodes and runs most efficiently when it has access to all of the hardware resources on each node.
So I wouldn't be surprised if they do have a total of 7500+ GPUs to balance workloads between. TO add, OpenAI has a long history of getting unlimited access to Google's clusters of GPUs (nowadays they pay for it, though). When they were training 'OpenAI Five' to play Dota 2 at the highest level, they were using 256 P100 GPUs on GCP[0] and they casually threw 256 GPUs at 'clip' for a short while in January of 2021[1].
As for how they do it, see these posts:
https://openai.com/blog/techniques-for-training-large-neural...
https://openai.com/blog/triton/