I understand that AMD GPUs offer better cost efficiency for F32 & F64 FLOPs, RAM, and wattage. But however, if ROCm is such a half baked piece, shouldn't that advantage be gone? What drives AMD adoption in the HPC space then?
I understand that AMD GPUs offer better cost efficiency for F32 & F64 FLOPs, RAM, and wattage. But however, if ROCm is such a half baked piece, shouldn't that advantage be gone? What drives AMD adoption in the HPC space then?
That’s been my experience with most large enterprise software, really. Amazing at its most critical competency, core dumps when I try something a little unusual.
More generally LLM training is weird because its a supercomputing workload being executed by Silicon Valley devs. So they want to use open source frameworks, find answers on stack overflow, use cloud providers etc. And it's new enough that random PHD's can invent stuff like Flash Attention without being employed by Nvidia.
Normal supercomputer users dont care about any of that.
Public bidding processes are different and much more reliant solely on price. As such you can expect them to pick different choices than the private market.
This can end up picking outright duds (see Aurora)
It's a good marketing metric, but probably contraproductive the AMDs longterm success in the field. They're spending engineering time building something they'll unlikely to be able to translate into other fields.
The DoE doesn't like having single suppliers for their tech. AMD/HPE's Frontier bid was a bit of a gamble from the DoE - it wasn't remotely obvious whether they'd be able to deliver it or not, and was against a background of Intel broadly failing to deliver Aurora - whereas nvidia was a known safe bet. However placing a big order with Intel and another one with AMD was their best shot at getting away from being wholly reliant on nvidia.
Frontier shipped. It's a real thing, people run code on it. I'd guess the DoE labs talk to each other to some extent and thus AMD ended up winning the El Capitan bid. That means AMD has a flagship HPC machine that sales people can point to and a big ongoing revenue stream to continue funding development from. It looks like the DoE plan to have two HPC vendors has worked.
Aurora seems to be somewhat in existence now but is less compelling as a story other potential customers might want to copy. I'm curious whether Intel end up making a loss on it.
So, as long as they can get a given set of applications to run above a given performance threshold, alternative vendors have a shot.