AMDs mastery of chiplet designs is in stark contrast to Nvidias reliance on extremely large and complex dies and pretty lack luster forays into multi-die in the past.
If Nvidia doesn't come up with ways to combat the inherent cost and yield advantages of chiplets they could get quite easily blown away on $/perf.
GPGPU didn't have enough money behind it to justify fully exploiting all hardware vs just going with whatever was easiest (Nvidia+CUDA) and eating the cost but that has changed substantially.
If you can buy AMD at a 10%+ $/perf advantage when you are talking about laying down $100M+ for accelerators on even a moderate sized cluster you can be damn sure the software will get tuned to take advantage of it.
It's the same TSMC tech that both Nvidia and AMD use, it's not something proprietary exclusive to one or another since they're both fables.
On the contrary, AMD has to be afraid since they have no answer to Nvidia's 800GB/s data center interconnects they got from their aquisition of Mellanox, which enables Nvidia to scale.
https://www.xilinx.com/products/technology/high-speed-serial...
Don't think that Infiniband is somehow an Nvidia moat, if anything because they ate up the last somewhat independent vendor (Mellanox) it's likely it will be replaced by some variant of what is now being called Ultra Ethernet which for the most part incorporates ideas from Infiniband into the Ethernet PHY layer along with bolting into libfabric at the userland level. The people with the brains behind it include the Omni-Path folk (which is descendent from QLogic and SilverStorm before that) which have extensive experience in Infiniband and are now free from the shackles of Intel.
However this has absolutely nothing to do with chiplet designs or multi-die architectures (which Nvidia has done before with their dual GPU stuff from the GTX 900 series days etc).
The fact of the matter is that Nvidia has no experience whatsoever with chiplets or advanced packaging. AMD has on the other hand absolutely tons from yes, chiplets but also HBM, 3d-vcache etc all of which are ultimately advanced packaging techniques that all build on shared background of expertise.
Nvidia needs to rapidly catch up in this area or their future is a lot less certain.
This is putting aside what is happening with Tenstorrent/Etched/Groq, etc which threaten to eat up the vast majority of the non-training workloads associated with the current crop of AI/LLM architectures.
> mostly built upon Python libraries that are platform agnostic, or at least can be. With enough support for AMD in Pytorch etc
You can't imagine how much and how long the work in "can be" is. Thus NVIDIA has two of the deepest moats you can have: time and money (you would probably have to outspend them by 10x for 10 years to catch up).
Source: go look at how the AMD backend is implemented in pytorch vs how the AMD backend is implemented.
Edit: people must really not understand that this code (ie building this stuff) is highly non-trivial. It's not like building a competing js frontend framework. You not only need good engineers you need lots of them because the surface area is very big - hardware platforms, especially AMD's aren't at all "abstract", so really to do it right you need a team per box in the product matrix.
NVidia has a massive headstart, but there are diminishing returns to their software work. If AMD performance per dollar is 10% better at the hardware level, and their software performance is 90% of that of nvidia they are ahead.
> I don't see it. Large customers are choosing AMD over nVidia.
Please show me term sheets of contracts signed with "large customers". Please include net, gross, and recurring revenue.
> diminishing returns to their software work
Makes zero sense: software penetration follows power laws (network effects etc) not exponential decay laws.
> If AMD performance per dollar is 10% better at the hardware level, and their software performance is 90% of that of nvidia they are ahead.
Again: you cannot fathom how enormous that "if" there is so I highly recommend you actually go look at the impl to get a sense.
Software libraries aren't social networks, though. According to this argument we should all still be using C and not seeing new languages take hold. There is almost certainly a point where a library stops being able to add new value by being extended and that competitors can catch up to relatively easily no matter what the user demographics look like.
There are definitely social aspects to software development so there might be some power laws to find; but building a competing ecosystem gets done regularly enough and works well (C# v. Java for example). Much easier to do than building a competing social network.
They literally literally are - I follow people on GH in order to check out what they're doing and what repos they're starring/using.
> Much easier to do than building a competing social network.
We're talking about the ecosystem around the base stack remember? So no it's not easier when the bulk of the work needs to come from OSS contributors.
People have the weirdest wishful thinking about this because they don't actually work in this area (and/or don't have stake in the game). If you're this confident please show me your portfolio allocation for NVDA and AMD. Otherwise hint: everyone working on this does not so casually dismiss NVIDIA's lead.
> We're talking about the ecosystem around the base stack remember? So no it's not easier when the bulk of the work needs to come from OSS contributors.
The ecosystem exists as an abstraction around a base stack though. Much like saying the existence of Linux solidifies x86 dominance - that isn't what happens in practice, the Linux distros just pick up support for whatever hardware API are floating around. At some point the base can't add more value by adding more API features, and then the middleware starts supporting multiple bases depending on hardware economics.
The OSS people have no incentive to lock themselves in to buying NVidia cards. If AMD cards tended to work for compute problems they'd probably much rather be using AMD. AMD has better support for anything OSS that isn't compute related.
Can you source that? That seems like a hard claim to source.
Everyone in the area agrees that they'd rather be using Nvidia hardware. I'd rather be using Nvidia hardware, indeed I switched to a Nvidia card. But that is not because I see Nvidia as having a long term defensible position, but because AMD's compute drivers on consumer cards appear to suck. It was implementation details and AMD's weaknesses, not Nvidia's strengths. Important details, but nothing that can be realistically called a moat and certainly not anything to do with CUDA itself. The issues were far more foundational.
Honestly it'd be interesting to have some real expert takes on what the problem with AMD is, because the major one I'm aware of was geohotz's and he didn't get much further than I did. Blockers came up much deeper in the stack than CUDA. ROCm would have been good enough except that it was running on AMD kernel drivers.
I don't know what relevance term sheets should have here, but AMD success at the HPC top end is pretty obvious if you look:
https://top500.org/lists/top500/2024/06/
AMD and nVidia are both powering about 1.5 Exaflop of compute of the top 10 supercomputer. Machines 1 and 5 are AMD, nVidia is number 3 and 6-10.
They are still massively behind nVidia in the market overall: They are projecting a tenth of the datacenter revenues of nVidia. But that's still 4 billion in sales.
https://www.amd.com/en/newsroom/press-releases/2024-5-21-amd...
I believe 38k definitely qualifies as large.
Maybe you misread me. I didn't mean to write that this was the default choice for large customers. Just that there are large customers for whom it makes sense to go AMD.
The thing is, from where I am working and planning, nvidias cuda advantage is not a thriving community around it that would be hard to replicate. The community aspects are much more prominent higher up the stack, if you support pytorch and tensorflow you have a ton of community. Nvidia absolutely rocks at having high performance proprietary libraries for every niche use case.
That's going to take time focus and investment, but the hand performance tuning of these libraries might only buy you so much over the pytorch version. Unless you are doing LLMs and pushing hard against the limits of what's possible, that last bit of tuning can wait.
Nvidia had a call with us not long ago. I genuinely don'tsee the moat. If we manage to launch a product in our space it will run on anything pytorch runs. There is no advantage to marrying ourselves to cuda.
I know - my group has several 100k hour allocations on frontier (had - I graduated in April). The difference between frontier and aurora and wherever buying GPUs and FAANG buying them is the labs don't refresh every 2 years. That's why I don't consider them "large customers". They're not even fullride customers - you can be sure frontier got a very nice deal, much nicer than FAANG would, because of the top500 prestige.
> The thing is, from where I am working and planning, nvidias cuda advantage is not a thriving community around it that would be hard to replicate. The community aspects are much more prominent higher up the stack, if you support pytorch and tensorflow you have a ton of community.
First sentence and the others are in direct contradiction. And literally it was the first thing I pointed out: every armchair quarterback thinks supporting pytorch and tensorflow is so easy but you can compare every how AMD is currently supported to know that it's not the case.
And anyway, my point was merely that there are large customers going AMD.
As for the supposed contradiction: Cuda is much more than PyTorch support. There is a real breadth of proprietary C++ libraries. When I (and many others here I guess) say cuda, we mean that breadth of effort. Not just the cuda backend of pytorch. You don't have to match that breadth of effort to be competitive for very many AI applications being built higher up the stack, you only have to have a good enough ROCm backend for pytorch, and from everything I am hearing, it's getting there (if you are on Instinct, and not consumer hardware). But I don't have first hand experience.
What I have first-hand experience with is what NVIDIA tried to pitch us. They want to be the AI "platform", rather than just a hardware vendor. It all sounded like they know that their software advantage is brittle in places, hence the "platform" strategy. But they couldn't really articulate what that should mean.
I think the moat metaphor is ill-suited. This is not a situation where, if the moat is breached in some places, the castle falls. NV will have a software advantage in some domains for a long time to come. But at the same time, we are seeing that AMDs hardware can be competitive and better in domains where the software advantage matters comparatively less.
Is your claim really that national labs outspend FAANG for compute? Like do you understand what you're saying? You're off by probably an order of magnitude if not two.
> You don't have to match that breadth of effort to be competitive for very many AI applications being built higher up the stack, you only have to have a good enough ROCm backend for pytorch, and from everything I am hearing, it's getting there (if you are on Instinct, and not consumer hardware). But I don't have first hand experience.
I don't want to say too much more because I work somewhere that is very close to this story/melodrama but everything you're hearing is aspirational hype and what I wish for everyone talking about this is much less gossip and much more first-hand reality.
Then you throw out vague statements, and claim that you "don't want to say much more", to complain in the next sentence that it's all gossip.
You have literally offered no argument, other than saying "trust me, I know, I am an insider"...
I absolutely did but seemingly no one understands the force of it (because no one actually cares about details). My argument was very very simple: go look at the hip backend in pytorch and the cuda backend. To anyone that understands what's what, that alone speaks volumes. Not my fault if you don't understand though.
The AMD and the CUDA backend in pytorch are actually derived from the same code base. AMD compiles the CUDA code to its HIP interface (which is essentially a reimplementation of CUDA), and that's that. It's all upstreamed and so changes to the CUDA backend are tested against the AMD build.
If there is a problem with that strategy, then it's up to you to demonstrate it because the prima facie evidence is that it works: I can go to Azure and buy an AMD Instinct VM and run a Hugging Face model on that right now, with marginally more effort than it takes to run it on an nVidia VM.
It's not ad hominem - it was my argument in literally my first response without a single word of shade. you then go and say "you have no argument" and I respond with "well if you can't understand the argument then sure I have no argument". pretty simple. I read this on reddit at some point: you're not the victim here, you're just starting a fight and then losing that fight.
> If there is a problem with that strategy, then it's up to you to demonstrate it because the prima facie evidence is that it works
you think that yolking yourself to your direct competitor's runtime API is a sound product strategy? really? i'm sure you'll come back with "intel x86_64 blah blah blah". the fact is it's not a sound strategy and maybe it would be wise for AMD to build out a real HIP backend (and maybe they already are).
You only said "look at the implementation" and said that if that's not obviously problematic to me, I just don't understand. If you had explained why you find it problematic we might have had a conversation on the viability of that strategy and the need for other complementary strategies.
Seriously, the only reason I am still here is because I hope you can understand why your replies were really not appropriate for facilitating interesting conversation.
To this day I prefer to work on my Fuel than any other machine. But it's still dead.
Again, I’m no expert on CPUs but I think you could have argued something similar about AMD vs Intel 12 years ago. Now it’s clear AMD’s tech is superior.
Haven't seen that land meaningfully so far, but it is an interesting risk to Nvidia.
CPU was easy, GPU was hard. The major difference I could spot though was all the time I was spending batting data around between different buffers, the APIs weren't really limiting me.
https://docs.nvidia.com/cuda/cuda-c-programming-guide/#maxim...
The library writers are obviously a lot better at all this than I am, but the experience of owning an AMD card was that they literally didn't and my experience left me believing that the reasons were more specific than "no CUDA". My experience was everything compute related caused crashes.
And with HIP, the user-facing memory model is virtually identical to CUDA to my knowledge. Converting a CUDA program to HIP, assuming no unsupported features or platform details are used, is basically just "do s/cu/hip/g".
Also, HIP wasn't around in 2014. To program an AMD GPU back then, you were probably using OpenCL, which is nicely platform independent but also arguably somewhat archaic.
I mean you're literally using the calvinball programming language - it's just a metafunctor in the language of macros, what's the big deal?
literally my (sanitized) code from 10 years ago
```cuda
__global__ void kernel_myKernel(int num_items, myTask_t * task_item,
#if VALIDATION == 1
myDebug1_t * output_debug1_arr, myDebug2_t * output_debug2_arr,
#endif myOutput_t * output_arr, int current_iteration)
```those myDebug_t pointers are pointers to host memory, so transferring the data back to host memory is transparently managed by CUDA or Thrust, and you just have an #IF or #IFDEF block in the code that dumps intermediate state to the pointer if validation is enabled in the build. And then you can run whatever test suite, on the actual intermediate data that's happening inside your kernel, so you can test your invariants/assertions. Test-suite lite edition/just the part you need - but you're testing actual kernel invocations, not just theoretical. And you can go to the level of dumping intermediate state of prefix_sums/reductions (or generalized "intermediate work items") every warp/grid iteration if you want - it's debug mode.
but granted this relies on the ability to have data dynamically piped between device and host memory spaces... which was a novel bit of syntactic sugar they added to the CUDA toolkit back in like 2012 (it was big around the Kepler era and it was a big feature on the Jetson TK1 too) so AMD might well have not had it. Or they might have said they had it but it just broke horribly if you used it.
but this is literally like three lines of code, if you had just used NVIDIA instead of AMD, because the feature was there when you needed it. Technically can be accomplished by just allocating extra VRAM for the buffer and copying it back afterwards, but granted, more work there.
What's the price of getting your work done? What's your time worth? That's always been the only moat of relevance that CUDA has had... and it's always been worth the expense.
Nvidia has multiple advantages.
1. Same software written in PyTorch works across all relatively modern Nvidia chips. With AMD and ROCm that's not the case.
2. M̶I̶3̶0̶0̶x̶ h̶a̶s̶ n̶o̶ F̶P̶8̶ s̶u̶p̶p̶o̶r̶t̶.
3. Comparisons against H100 are always like this:
8x AMD MI300X (192GB, 750W) GPU
8x H100 SXM5 (80GB, 700W) GPU
The fair comparison would be against 8x H100 NVL (188GB, <800W) GPU
And that they never do.4. H100 is 21 month old architecture. MI300X is 7 months. Nvidia moved into new architecture every year pace. AMD is generation behind and must step up the pace. B100 comes out this year.
AMD is getting closer, but don't except them to catch Nvidia in not time.
Also, many of these comparisons use vLLM for both setups, but for Nvidia you can and should use TensorRT-LLM which tends to do quite a bit better than vLLM at high loads.
Less than that, we paid for ours in January and received it in March. The first batch had problems and we had to send them back, which took another 3 weeks. So, let's consider the start date closer to April.
~3 months.
vLLM is not quite as performant but it's pretty close in production environments.
On the 4bit L3 70b test, they are using AWQ. GPTQ with Marlin kernels (which now gets repacked on the fly in vLLM) is much faster in our tests, and has much better optimizations amongst vLLM contributors.
So yeah, this is sample size of n=1, as are my own tests. But my experiences have been much closer to the L3 8b fp16 charts, so I'm going to presume the bigger delta comes down to that.
EDIT: here is one of our own recent internal tests where we observed Marlin performing 25-30% better than AWQ in vLLM - https://miro.medium.com/v2/resize:fit:1400/format:webp/1*F9i...
I don't have a chart handy, but in our tests, TRT was ca. 10% faster (but it's much more difficult to setup, you have to convert models, etc). LMDeploy we tested months ago, maybe it's improved now, but they were making fairly wild claims that we didn't observe.
Basically, you shouldn't publish a benchmark without providing all of the details. How long did test run? How were requests staggered? What were the input/output token sizes?
A lot of these 256 in / 256 out benchmarks are useless. If your system prompt is 4k tokens, prefill becomes a serious issue. How performant is your continuous batching? Can you do chunked prefill? Are you running prompt prefix caching (and if so, how performant is that)? It's a lot more complex than simply generating some tokens at various batch sizes.
A few things in the linked article that make my eyebrows raise. Notably, they claim to achieve 206 t/s at bs=1 on a single A100 80GB (2 TB/s bandwidth) in fp16. That model is 14.5GB in size, and best case would require aggregate memory bandwidth of almost 3 TB/s to achieve.
The BentML you linked to, also ran on an A100 80GB, and achieved ~650 t/s at bs=10 so roughly 65 t/s per stream. Granted this is for an 8B param model, not 7B, and it will of course be faster at bs=1 but absolutely not by this margin. In this same benchmark, they achieved roughly 45 t/s per stream at bs=50 so you can see the sort of scaling we're dealing with. At bs=10 you should achieve relatively similar per-stream throughput to bs=10.
Baseten did a benchmark with TensorRT-LLM and an A100 80GB on Mistral 7B and at bs=1 they achieved 75 t/s and at bs=8 they achieved 66 t/s per stream (https://www.baseten.co/blog/unlocking-the-full-power-of-nvid...)
I strongly suspect there is something very wrong with the Premai (never heard of them before) or there has been some huge inference breakthrough that I am unaware of. The BentoML benchmarks look pretty good, and I suspect there is some vLLM performance left on the table, however not enough to close the gap. In our testing, TensorRT-LLM was definitely faster, but not enough to warrant all of the other headaches.
“We are tight on supply, so there’s no questions in the near term that if we had more supply, we have demand for that product, and we're going to continue to work on those elements as we go through the year,” Su said, adding that she’s “very pleased with how the ramp is going.”
Nvidia has shipped almost 4 million DC GPUs in 2023 to give your an idea about the sizes for it to count on the balance sheet.
We are a brand new startup and tiny. Not pretending to be anything else. But we are also early and one of the first to buy and receive these cards.
We have excellent backing and will grow with demand. The more people wake up to an alternative, that is actually available and as we are proving… viable, the more demand we will have. The need for compute is endless.
Dell is a partner of ours and sees the value in our proposition. That is huge. They are hungry to grow their AI business the same way AMD is.
Give it time. The frankness is appreciated, but obvious. Comparisons against nvidia today are laughable. This is a long tail game. Talk to me in a few years.
This is not a football game with a winner and loser. Our goal is to not just be AMD, but every chip that our customers want.
Geohot's streams show you why that's a bad idea in real time.