I purposefully used the word feature in my reply and I don't think CUDA is a feature. I see that as a massive ecosystem built on a proprietary platform. For that, it takes orders of magnitude more money and something else money can't buy.. time. Time to have the platform adopted by countless vendors and the ecosystem built up.
GPUs are scarce, expensive and in high demand. Devs are cheap compared to the expense of training a huge model. If Cerebras or some other hardware company develops a viable competitor to the H100, it doesn't matter if the ISA is 36-bit VLIW documented in Vedic Sanskrit, they'll have infinite demand.
AMD has a viable competitor to H100 but no one buys it because it doesn’t support CUDA.
It’s actually better than the H100 by a mile in some inference workloads, but in others falls behind.
that's surprising to me. Torch supports ROCm, why are startups wasting their money on H100s if they can get a better deal with AMD GPUs?