You're asking the right question but I think the math is off by a bit. The equivalent number on the H100's is 989 TFLOP/s/chip so the equivalent job is ~10K H100's = (10 * 10^18) / (989 * 10^12). (Both chips also have 8-bit acceleration!)
I believe this is the largest ML job both by exaflops and number of chips every demonstrated. Other companies own more chips or exaflops than we show in this job but getting all the hardware working at once on a single job is a different matter! :-)
989 is TF32 core, for 16 bit it is 1979, so I guess around 5000 H100’s in a single training job would be equivalent to the training job mentioned in this article.
Either way I actually would not be surprised if OpenAI has launched a single job on more than 10k GPU’s, but I also am not very knowledgeable on practical scaling. Congrats on the feat!
I'm not certain but I think part of this is that XLA (for example) is a mountain of chip-specific optimizations between your code and the actual operations. So comparing your throughput between GPU and TPU is not just flops-to-flops.
The PaLM paper linked in the blog post is about how to get something actually useful out of that compute.
https://inflection.ai/inflection-ai-announces-1-3-billion-of...
As of 2 months ago, they had at least 7000 up and running, fwiw
https://youtu.be/z3hmfSVmyqg?feature=shared&t=3328
I'm curious why it is so hard for them to deploy the compute. They seem to be fairly behind schedule.
On what number or op for the h100?