https://www.nvidia.com/en-us/data-center/products/ai-enterpr...
It is this kind of delivery that the competition misses out.
https://www.nvidia.com/en-us/data-center/products/ai-enterpr...
It is this kind of delivery that the competition misses out.
The networking alone is a huge bottleneck at scale. A competitor has to be better at networking AND chips to be competitive.
NVIDIA got big because CUDA works on the most crappy notebook GPUs up to their most powerful chips, and AMD should do the same, but focusing their limited number of driver devs on the expensive enterprise hardware makes sense IMHO.
AI/ML is a rapidly moving field etc and you know geohot is gonna leak it all on twitter as soon as there's anything to announce, which makes it far more difficult for them to pivot later, etc.
Agree and yet none of the contenders were able to work out their software play (Intel, AMD, chip startups) for more than a year which shows how corporates move slow.
Google is not selling their TPUs AFAIK and their tooling is completely focused on internal use.
So really interesting to see no one else is properly addressing the need even though they have chips (and the chip itself is much simpler than a cpu, a systolic matrix multiplier array).
https://stability.ai/news/putting-the-ai-supercomputer-to-wo...
I’m sure this was a lot of work, and Intel surely helped a lot, and there are probably plenty of kludges involved. But it worked, and there’s a lot of money on the table to do things like this.
CUDA will do whatever you want and it more-or-less just works. ROCm (after > six years) is still:
- Won't work on your hardware
- Used to work on your hardware but we removed support within a few years
- Burn 10x more time trying to get something to work
- Be perpetually behind CUDA in terms of what you want/need to do
- Sorry, that just won't work
- Performance is lower than it should be for what is often actually better hardware, to the point where a superior newer generation AMD GPU gets bested by a previous generation Nvidia GPU with inferior (on paper) hardware specs
I've been trying ROCm since it was initially released > six years ago. I want AMD to succeed - I've purchased every new generation of AMD GPU in these six years to evaluate the suitability of AMD/ROCm for my workloads. Once a quarter or so I check back in to evaluate ROCm.
Every. Single. Time. I come away laughing/shaking my head at how abysmal it is. Then I go back to CUDA and sit in wonder at how well it actually works and throw even more money at Nvidia because I just need get things done and my concerns about their monopoly, artificial market segmentation, ridiculously high margins, etc are a distant second to my livelihood.
AMD (and others) need to understand what Jensen Huang has been saying for years - 30% of their development spend is on software. As the announcements this week show, Nvidia is using their greater and greater financial resources and market share to continue to lap AMD in the only thing people actually care about: here's our product and here's what you can actually do with it.
Many people with a fundamental hate/disgust for Nvidia will come back and say "ok bootlicker, it's supported in torch you're spreading FUD". Ok, take a look at the Nvidia platform you linked and show me where the ROCm equivalent is. Take a look at inference serving platforms which are one of the things I care most about. Look at flash attention, alibi, and the countless other software components that you actually need beyond torch in many cases. Watch even basic torch crash all over the place with ROCm.
Sure, you /might/ be able to train or run local one-off inference with AMD. How do I actually run this thing for my users? Crickets -or- maybe vLLM support for ROCm for LLMs (nothing for other models). Then dig just a little bit deeper and realize even vLLM isn't feature complete, requires patches, specific versions all around, and from personal experience a lot of github/blog spelunking and pain. With CUDA it's `docker run` and flies.
With CUDA I can run torchserve, HF TGI, vLLM, Triton, and a number of others to actually serve models up for users so I can make money from my work. ROCm, meanwhile, can barely run local experiments.
AMD needs to get it together.