Note: previous testing I did was on a single (8x) MI300X node, currently I'm doing testing on just a single MI300X GPU, so not quite apples-to-apples, multi-GPU/multi-node training is still a question mark, just a single data point.
Their MI300s already beat them, 400s coming soon.
Know any LLMs that are implemented in CUDA?
And here is pretty damning evidence that you're full of shit: https://github.com/ggml-org/llama.cpp/blob/master/ggml/src/g...
The ggml-hip backend references the ggml-cuda kernels. The "software is the same" (as in, it is CUDA) and yet AMD is still behind.
Show me one single CUDA kernel on Llama's source code.
(and that's a really easy one, if one knows a bit about it)
It is the same PyTorch whether it runs on an AMD or an NVIDIA GPU.
The exact same PyTorch, actually.
Are you're trying to suggest that the machine code that runs on the GPU is the one that is different?
If you knew a bit more, you would know that this is the case even between different generations of GPUs of the same vendor; making that argument completely absurd.