But in this case what is really ridiculous is the compute requirement. The required compute for optimal model growth roughly quadratically (both your model and your data grow linearly). So for 100T model you need 1e30 FLOPs. This machine gives you 1e18 FLOPs per second. It will take 30k years to train this model on one of these (or 30k of these to train it in a year, but then utilization will start kicking in).