Another piece of information is that CUDA software was provide free or cheaply to Universities doing LLM research I think. And the software is easy to use.
Another piece of information is that CUDA software was provide free or cheaply to Universities doing LLM research I think. And the software is easy to use.
Its competitors are only way behind when it comes to software support. The hardware coming out of Intel and amd is, especially for its price, very capable. Given how much money is being invested in AI right now, I don’t see Nvidia’s moat lasting more than a few more years.
Either you're the type of company that does that, or you aren't.
Getting good AI talent now is very costly. HW engineers are cheaper.
Nvidia has more SW than HW engineers for a reason and the transformation for that started slowly almost 2 decades ago and accelerated 2012 with AlexNet, the first public showcase of a NN running on GPUs. Jensen saw what that meant and transformed the company from that moment focusing on DeepLearning.
Nvidia isn't waiting for a market to develop but prefers to create markets by tackling hard and complex problems. It seems that Nvidia got lucky with AI but it was a long lasting preparation for Jensen.
Tell me though, what Fortune 500 do you know that is willing to put all their eggs in one basket? It is MBA 101 to not do that.
There needs to be alternatives in the space. Why not let them try?
I only dabble in AI stuff but have decades of experience doing quick surface-level quality checks of open source projects. I looked at some of AMD's ROCm repos late last year. Even basic stuff like the documentation for their RNG libraries didn't inspire confidence. READMEs had blatant typos in, everything gave off a feeling of immense lack of effort or care. Looking again today the ROCrand docs do seem improved, at least on the surface, I haven't tried it out for real.
But if we cast the net a little wider again, the same problems rear their ugly head. Flash Attention is a pretty important kernel to have if working with LLMs, maybe I'd like one of those for AMD hardware?
https://github.com/ROCm/flash-attention
We're in luck! An official AMD repo with flash attention in it, great! Except.... the README says at the top:
Requirements: CUDA 11.4 and above. We recommend the Pytorch container from Nvidia, which has all the required tools to install FlashAttention.
Really? Ah, if we scroll down all the way to the bottom we can find a new section that says "AMD/ROCm: Prerequisite: MI200 & MI300 GPUs". Guys, why not just rewrite the README, literally the first thing you see, to put the most important information up front? Why not ensure it makes sense? It takes 10 seconds and is the kind of attention to detail that makes me think the rest of your work will be high quality too.
Checking the issue tracker we see people reporting that the fork is very out of date, and that some models just mysteriously don't work with it due to bugs. These issue reports go unanswered for months. And let's not even go there on the hardware compatibility front, everyone already knows what "AMD support" really means (not the AMD cards you might actually own) vs what "NVIDIA support" means (any device that supports the needed CUDA version, of any size).
I would never try to defend AMD with regards to them needing to catch up. Even talking with executives at AMD, neither would they. Nobody is trying to pull a fast one on this.
What has changed for certain, is their attitude and attention. I just got back from Dell Tech World. Dell was caught off-guard with this AI thing too. It is obvious the only thing that anyone is talking about now is "ai ai ai ai ai ai".
Give them a bit of time and I think they will start to become competitive over the next few years. It won't happen over night. You won't see README's fixed right away. But one thing that is for certain, they are all at least trying now, instead of pretending it doesn't exist.
Whether they will be successful or not, is yet to be seen. I wouldn't even know how to define successful. I don't think anyone is kidding themselves about Nvidia being dominant. But, I'm personally willing to bet on them selling a lot of hardware and working on their software story.
You might not, and that is fine too.
Not only that, but it is all being done in the open, unlike their competition. Hotz demanded some documentation, they provided it and he still complained. Some people just can't find happiness.
Now, whether or not I am pushing them forward is yet to be seen, but at least I'm trying. By positioning myself as a new startup who's trying to help... that will easily garner all their support as well. As I said in another comment, why not let them try too?
First off, it’s a HW/SW solution and things like CUDA/NCCL/etc make a HUGE difference.
Second, the token/watt ratio of every other option is nearly an order of magnitude difference in real world tests. When you add in custom silicon like moronic Grok/Dojo and you see that there aren’t really any close competitors when using custom spins. That is money down the drain IMO. Best bet for most enterprises is to buy 25% AMD and 75% H100 if they can get it.
I think Blackwell is potentially a long term generational problem due to power limitations in most data centers for now.
If I can save 20% of my data center costs and cut a price-gouging vendor while bringing the solution in-house at a big tech org I am a hero.
Consumers won’t buy a Surface because Microsoft isn’t cool.
B2C will first ask about security and stability.
Do you think AWS, Azure and GCP are the cheapest cloud offerings? Of course not, but why do they dominate cloud computing in B2C while price gouging everyone?
Because they offer something beyond price and that is security and stability as well as a reliable partner. They also offer support and capacity on a level which a startup CSP will never be able to offer.
This is also the reason why all AI accelerator competitors won't be a competition for Nvidia.
To beat Nvidia it's not only about beating CUDA, it's about beating Nvidia Enterprise AI suite with it's security offerings and support options. But enterprise business level SW is a level where AMD and others will never go to and will have to rely on Big Tech like MS, Amazon and so on to do that for them. But why should they if they have in-house solutions? Big CSPs developing their own AI accelerators shows you that they understand Nvidia's business model and are trying to compete head on because they understand that Nvidia is attacking them at enterprise level with AI enterprise solutions. And of course any enterprise using Nvidia enterprise SW will automatically use Nvidia HW.
Once SW is more spread than HW then it dictates where the direction goes. If MS releases Windows 12 only for ARM then Intel and AMD are immediately screwed and they can't do nothing about that. No enterprise in the world cares if their CAD system runs on x86 or ARM as long as it can be used for the intended use.
If I am in charge of a data center I had better understand the impact of security and stability as well as the qualities of vendor relationships on my costs or I probably won’t be in that role very long.
You, on the other hand, apparently have never managed an enterprise ISA transition, or even cross-compiled software. The idea that Microsoft would just do that and that it would work is naive in the extreme. CAD software is compiled first for an architecture, and then generally within an operating system. It is all interconnected and interdependent.
I know that it's touted as the key competitive advantage, but it seems to stem from the fact it actually works, unlike others.
Still great advantage, but not a lock in. If competitors get their act together, couldn't they just replace CUDA with another API, all hidden somewhere in the sw stack?