Could someone explain why ROCm is so behind Cuda and why wouldnt big tech and startups adopt MI300X instead of waiting in line begging Jensen for H100?
Nvidia has been working on CUDA for 16 years and for the past few years building ML use cases on CUDA. AMD is late in the game with ROCm. That said, I think the future is bright for AMD. The MI300X looks great and they're spending a lot on making ROCm better and fully PyTorch compatible, which will automatically make all of the PyTorch-based software AMD ready.