NVIDIA got big because CUDA works on the most crappy notebook GPUs up to their most powerful chips, and AMD should do the same, but focusing their limited number of driver devs on the expensive enterprise hardware makes sense IMHO.
AI/ML is a rapidly moving field etc and you know geohot is gonna leak it all on twitter as soon as there's anything to announce, which makes it far more difficult for them to pivot later, etc.
The networking alone is a huge bottleneck at scale. A competitor has to be better at networking AND chips to be competitive.
Agree and yet none of the contenders were able to work out their software play (Intel, AMD, chip startups) for more than a year which shows how corporates move slow.
Google is not selling their TPUs AFAIK and their tooling is completely focused on internal use.
So really interesting to see no one else is properly addressing the need even though they have chips (and the chip itself is much simpler than a cpu, a systolic matrix multiplier array).
https://stability.ai/news/putting-the-ai-supercomputer-to-wo...
I’m sure this was a lot of work, and Intel surely helped a lot, and there are probably plenty of kludges involved. But it worked, and there’s a lot of money on the table to do things like this.