334 karma · joined February 23, 2019
> That sort of stuff causes pitchforks to rise up in other countries.
(Not that I agree)
Edit: In my opinion at least, maybe they would say that if models are exhibiting that stuff 20% of the time nowadays then we’re a few years away from that reaching > 50%, or some other argument that I would disagree with probably
* You can use less GPUs if you decrease batch size and increase number of steps, which would lead to a longer training time
* FP8 is pretty efficient, if Grok was trained with BF16 then LLama 4 should could need less GPUs because of that
* Depends also on size of the model and number of tokens used for training, unclear whether the total FLOPS for each model is the same
* MFU/Maximum Float Utilization can also vary depending on the setup, which also means that if you're use better kernels and/or better sharding you can reduce the number of GPUs needed
Is the new license different? Or is it still failing for the same issues pointed by the second point?
I think the problem with the 3rd point is that LeCun is not leading LLama, right? So this doesn't change things, thought mostly because it wasn't a good consideration before