I believe 38k definitely qualifies as large.
Maybe you misread me. I didn't mean to write that this was the default choice for large customers. Just that there are large customers for whom it makes sense to go AMD.
The thing is, from where I am working and planning, nvidias cuda advantage is not a thriving community around it that would be hard to replicate. The community aspects are much more prominent higher up the stack, if you support pytorch and tensorflow you have a ton of community. Nvidia absolutely rocks at having high performance proprietary libraries for every niche use case.
That's going to take time focus and investment, but the hand performance tuning of these libraries might only buy you so much over the pytorch version. Unless you are doing LLMs and pushing hard against the limits of what's possible, that last bit of tuning can wait.
Nvidia had a call with us not long ago. I genuinely don'tsee the moat. If we manage to launch a product in our space it will run on anything pytorch runs. There is no advantage to marrying ourselves to cuda.
I know - my group has several 100k hour allocations on frontier (had - I graduated in April). The difference between frontier and aurora and wherever buying GPUs and FAANG buying them is the labs don't refresh every 2 years. That's why I don't consider them "large customers". They're not even fullride customers - you can be sure frontier got a very nice deal, much nicer than FAANG would, because of the top500 prestige.
> The thing is, from where I am working and planning, nvidias cuda advantage is not a thriving community around it that would be hard to replicate. The community aspects are much more prominent higher up the stack, if you support pytorch and tensorflow you have a ton of community.
First sentence and the others are in direct contradiction. And literally it was the first thing I pointed out: every armchair quarterback thinks supporting pytorch and tensorflow is so easy but you can compare every how AMD is currently supported to know that it's not the case.
And anyway, my point was merely that there are large customers going AMD.
As for the supposed contradiction: Cuda is much more than PyTorch support. There is a real breadth of proprietary C++ libraries. When I (and many others here I guess) say cuda, we mean that breadth of effort. Not just the cuda backend of pytorch. You don't have to match that breadth of effort to be competitive for very many AI applications being built higher up the stack, you only have to have a good enough ROCm backend for pytorch, and from everything I am hearing, it's getting there (if you are on Instinct, and not consumer hardware). But I don't have first hand experience.
What I have first-hand experience with is what NVIDIA tried to pitch us. They want to be the AI "platform", rather than just a hardware vendor. It all sounded like they know that their software advantage is brittle in places, hence the "platform" strategy. But they couldn't really articulate what that should mean.
I think the moat metaphor is ill-suited. This is not a situation where, if the moat is breached in some places, the castle falls. NV will have a software advantage in some domains for a long time to come. But at the same time, we are seeing that AMDs hardware can be competitive and better in domains where the software advantage matters comparatively less.
Is your claim really that national labs outspend FAANG for compute? Like do you understand what you're saying? You're off by probably an order of magnitude if not two.
> You don't have to match that breadth of effort to be competitive for very many AI applications being built higher up the stack, you only have to have a good enough ROCm backend for pytorch, and from everything I am hearing, it's getting there (if you are on Instinct, and not consumer hardware). But I don't have first hand experience.
I don't want to say too much more because I work somewhere that is very close to this story/melodrama but everything you're hearing is aspirational hype and what I wish for everyone talking about this is much less gossip and much more first-hand reality.
Then you throw out vague statements, and claim that you "don't want to say much more", to complain in the next sentence that it's all gossip.
You have literally offered no argument, other than saying "trust me, I know, I am an insider"...
I absolutely did but seemingly no one understands the force of it (because no one actually cares about details). My argument was very very simple: go look at the hip backend in pytorch and the cuda backend. To anyone that understands what's what, that alone speaks volumes. Not my fault if you don't understand though.
The AMD and the CUDA backend in pytorch are actually derived from the same code base. AMD compiles the CUDA code to its HIP interface (which is essentially a reimplementation of CUDA), and that's that. It's all upstreamed and so changes to the CUDA backend are tested against the AMD build.
If there is a problem with that strategy, then it's up to you to demonstrate it because the prima facie evidence is that it works: I can go to Azure and buy an AMD Instinct VM and run a Hugging Face model on that right now, with marginally more effort than it takes to run it on an nVidia VM.
It's not ad hominem - it was my argument in literally my first response without a single word of shade. you then go and say "you have no argument" and I respond with "well if you can't understand the argument then sure I have no argument". pretty simple. I read this on reddit at some point: you're not the victim here, you're just starting a fight and then losing that fight.
> If there is a problem with that strategy, then it's up to you to demonstrate it because the prima facie evidence is that it works
you think that yolking yourself to your direct competitor's runtime API is a sound product strategy? really? i'm sure you'll come back with "intel x86_64 blah blah blah". the fact is it's not a sound strategy and maybe it would be wise for AMD to build out a real HIP backend (and maybe they already are).
You only said "look at the implementation" and said that if that's not obviously problematic to me, I just don't understand. If you had explained why you find it problematic we might have had a conversation on the viability of that strategy and the need for other complementary strategies.
Seriously, the only reason I am still here is because I hope you can understand why your replies were really not appropriate for facilitating interesting conversation.
https://www.amd.com/en/newsroom/press-releases/2024-5-21-amd...