NVidia has a massive headstart, but there are diminishing returns to their software work. If AMD performance per dollar is 10% better at the hardware level, and their software performance is 90% of that of nvidia they are ahead.
NVidia has a massive headstart, but there are diminishing returns to their software work. If AMD performance per dollar is 10% better at the hardware level, and their software performance is 90% of that of nvidia they are ahead.
> I don't see it. Large customers are choosing AMD over nVidia.
Please show me term sheets of contracts signed with "large customers". Please include net, gross, and recurring revenue.
> diminishing returns to their software work
Makes zero sense: software penetration follows power laws (network effects etc) not exponential decay laws.
> If AMD performance per dollar is 10% better at the hardware level, and their software performance is 90% of that of nvidia they are ahead.
Again: you cannot fathom how enormous that "if" there is so I highly recommend you actually go look at the impl to get a sense.
Software libraries aren't social networks, though. According to this argument we should all still be using C and not seeing new languages take hold. There is almost certainly a point where a library stops being able to add new value by being extended and that competitors can catch up to relatively easily no matter what the user demographics look like.
There are definitely social aspects to software development so there might be some power laws to find; but building a competing ecosystem gets done regularly enough and works well (C# v. Java for example). Much easier to do than building a competing social network.
They literally literally are - I follow people on GH in order to check out what they're doing and what repos they're starring/using.
> Much easier to do than building a competing social network.
We're talking about the ecosystem around the base stack remember? So no it's not easier when the bulk of the work needs to come from OSS contributors.
People have the weirdest wishful thinking about this because they don't actually work in this area (and/or don't have stake in the game). If you're this confident please show me your portfolio allocation for NVDA and AMD. Otherwise hint: everyone working on this does not so casually dismiss NVIDIA's lead.
> We're talking about the ecosystem around the base stack remember? So no it's not easier when the bulk of the work needs to come from OSS contributors.
The ecosystem exists as an abstraction around a base stack though. Much like saying the existence of Linux solidifies x86 dominance - that isn't what happens in practice, the Linux distros just pick up support for whatever hardware API are floating around. At some point the base can't add more value by adding more API features, and then the middleware starts supporting multiple bases depending on hardware economics.
The OSS people have no incentive to lock themselves in to buying NVidia cards. If AMD cards tended to work for compute problems they'd probably much rather be using AMD. AMD has better support for anything OSS that isn't compute related.
Can you source that? That seems like a hard claim to source.
Everyone in the area agrees that they'd rather be using Nvidia hardware. I'd rather be using Nvidia hardware, indeed I switched to a Nvidia card. But that is not because I see Nvidia as having a long term defensible position, but because AMD's compute drivers on consumer cards appear to suck. It was implementation details and AMD's weaknesses, not Nvidia's strengths. Important details, but nothing that can be realistically called a moat and certainly not anything to do with CUDA itself. The issues were far more foundational.
Honestly it'd be interesting to have some real expert takes on what the problem with AMD is, because the major one I'm aware of was geohotz's and he didn't get much further than I did. Blockers came up much deeper in the stack than CUDA. ROCm would have been good enough except that it was running on AMD kernel drivers.
I don't know what relevance term sheets should have here, but AMD success at the HPC top end is pretty obvious if you look:
https://top500.org/lists/top500/2024/06/
AMD and nVidia are both powering about 1.5 Exaflop of compute of the top 10 supercomputer. Machines 1 and 5 are AMD, nVidia is number 3 and 6-10.
They are still massively behind nVidia in the market overall: They are projecting a tenth of the datacenter revenues of nVidia. But that's still 4 billion in sales.
https://www.amd.com/en/newsroom/press-releases/2024-5-21-amd...
I believe 38k definitely qualifies as large.
Maybe you misread me. I didn't mean to write that this was the default choice for large customers. Just that there are large customers for whom it makes sense to go AMD.
The thing is, from where I am working and planning, nvidias cuda advantage is not a thriving community around it that would be hard to replicate. The community aspects are much more prominent higher up the stack, if you support pytorch and tensorflow you have a ton of community. Nvidia absolutely rocks at having high performance proprietary libraries for every niche use case.
That's going to take time focus and investment, but the hand performance tuning of these libraries might only buy you so much over the pytorch version. Unless you are doing LLMs and pushing hard against the limits of what's possible, that last bit of tuning can wait.
Nvidia had a call with us not long ago. I genuinely don'tsee the moat. If we manage to launch a product in our space it will run on anything pytorch runs. There is no advantage to marrying ourselves to cuda.
I know - my group has several 100k hour allocations on frontier (had - I graduated in April). The difference between frontier and aurora and wherever buying GPUs and FAANG buying them is the labs don't refresh every 2 years. That's why I don't consider them "large customers". They're not even fullride customers - you can be sure frontier got a very nice deal, much nicer than FAANG would, because of the top500 prestige.
> The thing is, from where I am working and planning, nvidias cuda advantage is not a thriving community around it that would be hard to replicate. The community aspects are much more prominent higher up the stack, if you support pytorch and tensorflow you have a ton of community.
First sentence and the others are in direct contradiction. And literally it was the first thing I pointed out: every armchair quarterback thinks supporting pytorch and tensorflow is so easy but you can compare every how AMD is currently supported to know that it's not the case.
And anyway, my point was merely that there are large customers going AMD.
As for the supposed contradiction: Cuda is much more than PyTorch support. There is a real breadth of proprietary C++ libraries. When I (and many others here I guess) say cuda, we mean that breadth of effort. Not just the cuda backend of pytorch. You don't have to match that breadth of effort to be competitive for very many AI applications being built higher up the stack, you only have to have a good enough ROCm backend for pytorch, and from everything I am hearing, it's getting there (if you are on Instinct, and not consumer hardware). But I don't have first hand experience.
What I have first-hand experience with is what NVIDIA tried to pitch us. They want to be the AI "platform", rather than just a hardware vendor. It all sounded like they know that their software advantage is brittle in places, hence the "platform" strategy. But they couldn't really articulate what that should mean.
I think the moat metaphor is ill-suited. This is not a situation where, if the moat is breached in some places, the castle falls. NV will have a software advantage in some domains for a long time to come. But at the same time, we are seeing that AMDs hardware can be competitive and better in domains where the software advantage matters comparatively less.
Is your claim really that national labs outspend FAANG for compute? Like do you understand what you're saying? You're off by probably an order of magnitude if not two.
> You don't have to match that breadth of effort to be competitive for very many AI applications being built higher up the stack, you only have to have a good enough ROCm backend for pytorch, and from everything I am hearing, it's getting there (if you are on Instinct, and not consumer hardware). But I don't have first hand experience.
I don't want to say too much more because I work somewhere that is very close to this story/melodrama but everything you're hearing is aspirational hype and what I wish for everyone talking about this is much less gossip and much more first-hand reality.
Then you throw out vague statements, and claim that you "don't want to say much more", to complain in the next sentence that it's all gossip.
You have literally offered no argument, other than saying "trust me, I know, I am an insider"...
I absolutely did but seemingly no one understands the force of it (because no one actually cares about details). My argument was very very simple: go look at the hip backend in pytorch and the cuda backend. To anyone that understands what's what, that alone speaks volumes. Not my fault if you don't understand though.
The AMD and the CUDA backend in pytorch are actually derived from the same code base. AMD compiles the CUDA code to its HIP interface (which is essentially a reimplementation of CUDA), and that's that. It's all upstreamed and so changes to the CUDA backend are tested against the AMD build.
If there is a problem with that strategy, then it's up to you to demonstrate it because the prima facie evidence is that it works: I can go to Azure and buy an AMD Instinct VM and run a Hugging Face model on that right now, with marginally more effort than it takes to run it on an nVidia VM.
It's not ad hominem - it was my argument in literally my first response without a single word of shade. you then go and say "you have no argument" and I respond with "well if you can't understand the argument then sure I have no argument". pretty simple. I read this on reddit at some point: you're not the victim here, you're just starting a fight and then losing that fight.
> If there is a problem with that strategy, then it's up to you to demonstrate it because the prima facie evidence is that it works
you think that yolking yourself to your direct competitor's runtime API is a sound product strategy? really? i'm sure you'll come back with "intel x86_64 blah blah blah". the fact is it's not a sound strategy and maybe it would be wise for AMD to build out a real HIP backend (and maybe they already are).
You only said "look at the implementation" and said that if that's not obviously problematic to me, I just don't understand. If you had explained why you find it problematic we might have had a conversation on the viability of that strategy and the need for other complementary strategies.
Seriously, the only reason I am still here is because I hope you can understand why your replies were really not appropriate for facilitating interesting conversation.