CUDA Moat Still Alive
semianalysis.com
semianalysis.com
This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.
Cutlass is a fine piece of engineering, but it is not quite as good as their closed source libraries in real world workloads. There is secret sauce that is not open sourced.
Spend a billion on AMD shares, Spend another Billion on a out-of-house software team to solve the software solution to more than double the share price.
Taking into account that there are players that already own billions in AMD shares, they could probably do that as well. On the other hand perhaps it would be better for them, as major shareholders, to have a word with AMD management.
Hopefully I am reading too much into this. Hopefully she doesn't have any weird hangups over investing in software and it all just takes time to Do It Right after GPGPU got starved in the AMD winter. But if it is a weird hangup then yeah, 100%, ownership needs to get management in line because whiffing a matmul benchmark years into a world where matmul is worth trillions just ain't it.
It's not a deflection, but a straightforward description of AMDs current top-down market strategy of partnering with big players instead of doubling down to have a great OOBE for consumers & others who don't order GPUs by the pallet. It's an honest reflection if their current core competencies, and the opportunity presented by Nvidia's margins.
They are going for a bang-for-buck right now aiming at data center workloads, and the hyperscalers care a lot about perf/$ than raw performance at. Hyperscalers are also more self-sufficient at software: they have entire teams working on PyTorch, Jax, and writing kernels.
I've been on both sides of this shitshow, I've even said those lines before! But I've also been in the trenches making the broken shit work and I know that it's fundamentally an excuse. There's a reason why people pay 80% margin to Nvidia and there's a reason why AMD is worth less than the rounding error when people call NVDA a 3 trillion dollar company.
It's not because people can't read a spec sheet, it's because people want their expensive engineers training models not changing diapers on incontinent equipment.
I hope AMD pulls through but denial is _not_ the move.
Would you say AMD is "shitting the bed" by not building it's own consoles too? You know AMD could build a kick-ass console since they are doing the heavy-lifting for the Playstation, and the XBox[1] , but AMD knows as much as anybody that they don't have the skills to wrangle studio relationships or figure out which games to finance. Instead, they lean hard in their HW skills and get Sony Entertainment/the Xbox division do what they do best.
1.and the Steam Deck, plus half a dozen Deck clones.
This is a case-specific example of failure, it doesn't generalise very well to other markets. AMD is really well positioned for this very specific opportunity of historic proportions and the only thing holding them back is a somewhat continuous stream of unforced failures when writing a high quality compute driver. It seems to be pretty close to one single team of people holding the company back although organisational issues tend to stem from a level or two higher than the team. This could be the most visible case of value destruction by a public company we'll see in our lifetimes.
Optimistically speaking maybe they've already found and sacked the individual responsible and we're just waiting for improvement. I'm buying Nvidia until that proves to be so.
Even Steam Deck is only a success, because it depends on Windows ecosystem, and the moment Microsoft decides it is enough, lets see how long it holds.
All the games that matter are Windows games running via Proton, as Valve has failed to actually build a GNU/Linux native games ecosystem, in spite of UNIX/POSIX underpinnings of Android NDK, PlayStation, the studios hardly bother.
The day Microsoft actually decides to challenge Proton, or do a netbooks move on handhelds with XBox OS/Windows, the SteamDeck will lose, just like the netboooks did.
Additionally, it is anyone's guess what will happen to Valve when Gabe steps down.
Everyone is quite curious what Microsoft will drop at CES 2025, and which OEMs will be on their side, it is going to be netbooks all over again.
None of this matters because AMD drivers are broken. No one is asking AMD to write a PyTorch backend. The idea that AMD will have twice the silicon performance than nvidia to make up the performance loss for bad software is a pipedream.
Do you honestly think the MI300 has show-stopper driver bugs, or that Meta/Amazon doesn't have a direct line to AMD engineers?
Yes
>Meta/Amazon doesn't have a direct line to AMD engineers?
I don't even think AMD engineers have a direct line to AMD.
How do you know that the problems arise from broken drivers rather than broken hardware? Real world GPU drivers are full of workarounds for hardware bugs.
AMD has to get on top of their software quality issues if they're ever going to succeed in this segment, or they need to be producing chips so much faster than Nvidia that it's worth the extra time investment and pain.
[citation needed]
The same article also states that AMD provided custom bug-fixes written by Principle Engineers to address bugs in a benchmark - this is software that will only become part of the public release in 2 quarters. I ask again, do you think AMD will not expedite non-public bug-fixes for hyperscalers?
> You can't train with AMD, full stop, because their software stack is so buggy.
Point 7 from the article:
>> The MI300X has a lower total cost of ownership (TCO) compared to the H100/H200, but training performance per TCO is worse on the MI300X on public stable releases of AMD software. This changes if one uses custom development builds of AMD software.
I have an inkling that Meta does not obtain MI300 drivers from https://download.amd.com
Which is why, as they say clearly, nobody is training models on AMD. Only inference, at most. I'm not sure why you keep claiming they are training using private drivers. They clearly aren't.
They have spare parts you'd bet, and I'd bet they have some SLA agreement with each customer where an engineer is basically on call nearby in case a single thing dosnt work or a random part breaks or needs servicing.
Asianometry did a great video on the cost of downtime when it comes to ASML device in any fab. While I am not directly in this field and can't speak to the accuracy of the numbers john gives, he does not seem one to just make stuff up as his quality of video production for niche topics is quite good.
Probably for the best though, KFAB had been discharging several tons of solvents, cleaning agents, and reagents per year into the surrounding area [for as long as it ran](https://enviro.epa.gov/facts/tri/ef-facilities/#/Release/640...)
https://enviro.epa.gov/facts/tri/ef-facilities/#/Release/640...
its nice to be aware of this but this is so fastly different from a critisism point of view that i don't think that matters.
The prize is trillions of dollars, and they can print hundreds of millions if they can convince the market that they are closing the gap.
It’s embarrassing that whoever actually tries to use their product hits these crass bugs (same with geohot who was really invested in making AMD’s cards work; I think he just ran their demo script in a loop and produced crashes).
It seems they really don’t understand/value the developer flywheel.
1) Cash is king
2) Inventory is evil
I think this mindset may still be here, in 2024
Infiniband is an industry standard. It is weird to see the industry invent yet another standard to do effectively the same thing just because Nvidia is using it. This “Nvidia does things this way so let’s do it differently” mentality is hurting AMD:
* Nvidia has a unified architecture so let’s split ours into RDNA and CDNA.
* Nvidia has a unified driver, so let’s make a different driver for every platform.
* Nvidia made a virtual ISA (PTX) for backward compatibility. Let’s avoid that.
* Nvidia is implementing tensor cores. Let’s avoid those on RDNA. Then implement them on CDNA and call them matrix cores.
* Nvidia is using Infiniband like the rest of the HPC community. Let’s use Ethernet.
I am sure people can find more examples. Also, they seem to have realized their mistake in splitting their architecture into RDNA and CDNA, since they are introducing UDNA in the future to unify them like Nvidia does.Ultra Ethernet is a joint project between dozens of companies organized under the Linux Foundation.
https://www.phoronix.com/news/Ultra-Ethernet-Consortium
>> The Linux Foundation has established the Ultra Ethernet Consortium "UED" as an industry-wide effort founded by AMD, Arista, Broadcom, Cisco, Eviden, HPE, Intel, Meta, and Microsoft for designing a new Ethernet-based communication stack architecture for high performance networking.
You probably can't call it "industry standard" yet but the goal is obviously for it to become one.
> It is weird to see the industry invent yet another standard to do effectively the same thing just because Nvidia is using it.
This is a misstep for all involved, AMD included. Even if AMD is following everyone else by jumping off a bridge, AMD is still jumping too.
Infiniband is not fun - it's a special snowflake of an interconnect that sits parallel to the rest of your datacenter network, and can not really run a standard TCP/IP codebase (yeah IPoIB is a thing but still). Do the Nvidia boxes really need a scale-out IP network as well as an Infiniband network?
Plus the spec is old. Packet spraying and trimming, better ordering guarantees, queue pair scalability... a whole bunch of enhancements have been incorporated into UE all the while being compatible with regular Ethernet.
Qlogic was never really an Infiniband vendor --- their qib driver is still in the Linux codebase and essentially emulates verbs on top of a messaging-based design.
No need for Nvidia to go first to an industry standard and neither for AMD.
Personally would be great its getting backported but its so far away from an normal use case.
https://en.wikipedia.org/wiki/InfiniBand
Infiniband is extremely popular in the HPC space, which is why Nvidia adopted it. Everyone else saw Nvidia adopt it and said "Let us make a new network standard to be incompatible". This is mind boggling.
Even more mind boggling is that many of the companies in the Ultra Ethernet Consortium are members of the Infiniband Trade Association, AMD included:
https://www.infinibandta.org/member-listing/
This would be like the automotive industry forming a consortium to invent new incompatible wheels to exclude a successful upstart that adopted their existing standard wheel designs. With trillions of dollars in revenue on the line, you would think that companies would use existing networking standards to focus on building competitive hardware with reduced time to market, yet they are instead reinventing networking standards just because they can. This is a huge gift to Nvidia, since it means that everyone else is wasting time and money instead of being competitive.
> Infiniband is extremely popular in the HPC space
Not anymore. There used to be Cray Aries/GNI, psm/psm2, and now there's Slingshot, the new Cornelis stuff etc. There's almost no Infiniband now.
https://www.infinibandta.org/infiniband-and-roce-advances-fu...
Where are you getting your information?
If you buy a single DGX H100 rack and run LINPACK, you automatically get TOP500-grade numbers. Infiniband is a solid product, if not the best commercial offering for AI/ML, but no one buys it for an HPC cluster separately from the DGX boxes.
https://www.top500.org/system/180171/
You can likely find more. Infiniband has been excellent for HPC since the 2000s. That includes all HPC workloads, not just AI/ML.
Excuse me if I do not believe your claims concerning infiniband. They contradict not only actual data, but also what I have heard from people I consider experts.
Also, you did not answer my question concerning the origin of your information. I notice from another comment if yours that you have been talking to a LLM about this conversation. Have you been posting things that a LLM tells you?
Infiniband is not an industry standard lol.
Maybe it used to be, but it definitely is not anymore. Most Infiniband vendors are dead. The only product from those days that endures is Cornelis' Omnipath, and even that only emulated the Infiniband API back with its first gen, and then evolved to be its own thing.
At this point, Infiniband is as good as a proprietary interconnect only sold by Nvidia/Mellanox.
https://www.infinibandta.org/member-listing/
As far as I know, everyone is free to sign up with the infiniband trade association and implement the specification.
Furthermore, if you use RDMA over Ethernet, you are using infiniband at a low level. RoCE which enables it was originally called Ethernet over infiniband. It is maintained by the infiniband trade association.
Omnipath was Intel’s failed effort to try to kill an open standard. It purchased QLogic’s infiniband business, killed it in favor of omnipath and sold it when it failed.
This is wrong.
RDMA over Ethernet is... RDMA over Ethernet. There is no Infiniband involved.
RoCE was motivated by supporting RDMA, which was then an IB-only feature, over regular Ethernet. The user-level APIs are the same (verbs), but the underlying architecture is all different --- it is traditional Ethernet with link-level flow control to make it lossless (pause frames).
(I did that yesterday and was very satisfied :) )
I am not interested in continuing this discussion, but if you want to do your own research, I suggest starting with the fact that the Infiniband Trade Association controls the RoCE specification. I suggest you avoid using LLMs, for obvious reasons.
I would at least hope that they know where the speed is going, but the issue of torch.matmul and F.Linear using different libraries with different performance suggests that they don't even know which code they are running, let alone where the slow bits in that code are.
It's easy enough that there's blog articles showing single developers getting within spitting distance of NVIDIA's highly optimised code. As in, 80-something-percent of the best available algorithms!
All NVIDIA did was "put the effort in", where the effort isn't some super clever algorithm implemented by a unique genius, but they simply made hundreds of variants of the matmul algorithm optimised for various scenarios. It's a kind of algorithmic brute force for eking out every last percentage point for every shape and size of input matrices on every GPU model and even for various SLI configurations.
From what I've seen, AMD has done... none of this.
There are a number of pull-requests to ROCMblas for tuning various sizes of GEMV and GEMM operations. For example: https://github.com/ROCm/rocBLAS/pull/1532
That’s about half a decade after they should have done this foundational work!
I guess it’s better late than never, but in this case a timely implementation was worth about a trillion dollars… maybe two.
https://salykova.github.io/matmul-cpu
Concidentally, the Intel MKL also outperforms OpenBLAS, so there being room for improvement is well known. That said, I have a GEMV implementation that outperforms both the Intel MKL and OpenBLAS in my tests on Zen 3:
https://github.com/ryao/llama3.c/blob/master/run.c#L429
That is unless you shoehorn GEMV into the Intel MKL's batched GEMM function, which then outperforms it when there is locality. Of course, when there is no locality, my code runs faster.
I suspect if/when this reaches the established amd64 BLAS implementations' authors, they will adopt my trick to get their non-batched GEMV implementations to run fast too. In particular, I am calculating the dot products for 8 rows in parallel followed by 8 parallel horizontal additions. I have not seen the 8 parallel horizontal addition technique mentioned anywhere, so I might be the first to have done it.
While there are complex state-of-the-art algorithms, those algorithms exist for everyone. The overhead is the bit that had to be done to make the algorithm work.
For instance for sorting a list of strings the algorithm might be quick sort. The overhead would be in the efficiency of your string compare.
For matmul I'm not sure what your overhead is beyond moving memory, multiplying, and adding. A platform touting a memory bandwidth and raw compute advantage should have that covered. Where is the performance being lost?
I guess the only real options are stalls, unnecessary copies, or unnecessary computations.
The use of the word 'algorithm' is incorrect.
Look... I do this sort of work for a living. There has been no useful significant change to matmul algorithms.
What has changed is the matmul process.
Modern perf optimization on GPUs has little to do with algorithms and everything to do with process optimization. This is akin to factory floor planning and such. You have to make sure the data is there when the processing units need it, and the data is coming in at the fastest rate possible, while keeping everything synchronized to avoid wrong results or deadlocks.
Really compute power has nothing to do with it. It's a waste of time to even consider it. We can compute matmuls much faster than you can naively bring memory to the processing units. Whoever solves that problem will become very rich.
To that end, NVIDIA ships libraries that will choose from a wide variety of implementations the appropriate trade-offs necessary for SoTA perf on matmuls of all shapes and data types.
So what is done in practice is an algorithm that doesn't calculate the same result, but is imperceptibly close to doing classic attention, namely flash attention. Flash attention lets you fuse the kernel so that you can multiply against the V matrix and therefore write the condensed output to HBM. As an additional benefit you also go from quadratic memory usage to linear memory usage. But here is the problem: Your SRAM is limited and even flash attention is still O(n^2) in compute. This means if you tile your K and V cache into j and k tiles. You will have to load j*k times from memory. Meanwhile compute tends to consume very little silicon area. So you end up in a situation where you have excessive compute vs your SRAM. In the compute > SRAM regime, doubling SRAM size also doubles performance. You're memory bound again for super long contexts.
Now let's assume the opposite. Your compute resources are improperly sized with regards to your SRAM, you have too much SRAM but not enough compute resources e.g. a CPU. You will be compute bound with regard to a linear factor vs your SRAM, but always memory bound vs main memory. You could add the matrix cores to the CPU and the problem would disappear in thin air.
https://github.com/ryao/llama3.c
The only thing you wrote that makes any sense to me is “Flash attention lets you fuse the kernel”. Everything else you wrote makes no sense to me. For what it is worth, flash attention does not apply to llama 3 inference as far as I can tell.
Isn't the fastest theoretical algorithm something like O(n^2.37) ?
What?
I forgot the log(log(n)) factor.
In any case, for matrix multiplications that people actually do, this algorithm runs slower than a well optimized O(n^3) matrix multiplication implementation because the constant factor in the Big O notation is orders of magnitude larger.
Good CPUs and GPUs have a throughput in Flop/s for matrix multiplication that is between 60% and 90% of the maximum possible throughput, with many (especially the CPUs) reaching values towards the high end of that range.
As shown in the article, the AMD GPUs attain only slightly less than 50% (for BF16; for FP8 the AMD efficiency is even less than 40%).
Such a low efficiency for the most important operation is not acceptable.
I remember geohot saying something similar about a year ago
"We recommend that AMD to fix their GEMM libraries’ heuristic model such that it picks the correct algorithm out of the box instead of wasting the end user’s time doing tuning on their end." Is such a profoundly unhelpful thing to say unless you imagine AMDs engineers just sitting around wondering what to do all day.
AMD needs to make their drivers better, and they have. Shit just takes time.
If there are bugs in AMD code that prevent running tests, I bet there are even more bugs that don't manifest until you look at results.
I still think it is a mistake to say that CUDA is a moat. IMO the problem here is that AMD still doesn't seem to think that GPGPU compute is a thing. They don't seem to understand the idea that someone might want to use their graphics cards to multiply matricies independently of a graphics pipeline. All the features CUDA supports are irrelevant compared to the fact that AMD can't handle GEMM performantly out of the box. In my experience it just can't do it, back in the day my attempts to multiply matrices would crash drivers. That isn't a moat, but it certainly is something spectacular.
If they could manage an engineering process that delivered good GEMM performance then the other stuff can probably get handled. But without it there really is a question of what these cards are for.
GPU support lagged behind for years, no support for APUs and no guaranteed forward compatibility were clear signs that as a whole they have no idea what they are doing when it comes to building and shipping a software ecosystem.
To that you can add the long history of both AMD and ATI before they merged releasing dog shit software and then dropping support for it.
On the other hand you can take any CUDA binary even one that dates back to the original Tesla and run it on any modern NVIDIA GPU.
This particular difference stems the fact that NVIDIA has PTX and AMD does not have any such thing. Ie this kind of backwards compatibility will never be possible on AMD.
Having to create a binary that targets a very specific set of hardware and having no guarantees and in fact having a guarantee that it won’t on future hardware is what make ROCM unusable for anything you intend to ship.
What’s worse is that they also drop support for their GPUs faster than Leo drops support for his girlfriends once they reach 25…
So not only that you have to recompile there is no guarantee that your code would work with future versions of ROCM or that future versions of ROCM could still produce binaries which are compatible with your older hardware.
Like how is this not the first design goal to address when you are building a CUDA competitor I don’t fucking know.
The words "tech debt" do not have any meaning at AMD. No one understands why this is a problem.
This is likely self inflicted. They decided to make two different architectures. One is CDNA for HPC and the other is RDNA for graphics. They are reportedly going to rectify this with UDNA in the future. However, that is what they really should have done from the start. Nvidia builds 1 architecture with different chips based on it to accommodate everything and code written for one easily works on another as it is the same architecture. This is before even considering that they have PTX to be an intermediate language that serves a similar purpose to Java byte code in allowing write once, run anywhere.
They didn’t release support even for all GPUs from the same generation and dropped support for GPUs sometime within 6 months of releasing a version that actually “worked”.
The entire core architecture behind ROCM is rotten.
P.S. NVIDIA usually has multiple CUDA feature levels even within a generation. The difference is that a) they always provide a fallback option, and usually this doesn’t require any manual intervention and b) is that as long as you define the minimum target framework when you build the binary you are guaranteed to run on all past hardware that is supported by the feature level you targeted and on all future hardware.
https://docs.nvidia.com/cuda/parallel-thread-execution/index...
They also appear to be cululative.
You don’t get that with ROCm, and this is why it’s garbage unless someone else abstracts all of that from you.
So if Microsoft is happy to maintain an ML as a service solution that just takes prompts and maybe data it’s not your problem.
But if you need to run your own workloads and these can include workloads that are well outside of “AI” and might not be even possible or remotely profitable to have a SAAS wrapper around them it’s all on you.
I'm particularly interested in building a Home Assistant machine that can run the voice assistant locally (STT/TTS/LLM) while using the least amount of power / generating the least amount of heat and noise.
What they write about ROCm and Windows is equivocation. They target only one app: Blender. Pytorch+ROCm+Windows does not work.
I had bought a 6900XT myself around launch time (the RTX3080 I ordered was not coming, it was the chip shortage times...) and it took around 2 years for Pytorch to become actually usable on it.
I remember CUDA being much more buggy back then but it still worked pretty good.
Back then AMD wasn't considered a real competition for ML/AI hardware.
Glad as always to see more competition in the market to drive innovations. AMD seems to be letting larger VRAM onto consumer cards, which is nice to see, just hope the AI/ML experience can get better for their software ecosystem.
That said, I would not expect it to stay working for long as long as ROCm is a dependency since AMD drops support for its older GPUs quickly while Nvidia continues to support older GPUs with less frequent legacy driver updates.
Did make me wish I bought a Nvidia.
https://ameridroid.com/products/home-assistant-voice-preview...
"Met with @LisaSu today for 1.5 hours as we went through everything
She acknowledged the gaps in AMD software stack
She took our specific recommendations seriously
She asked her team and us a lot of questions
Many changes are in flight already!
Excited to see improvements coming"
"Thanks @dylan522p for the constructive conversation today. Feedback is a gift even when it’s critical. We have put a ton of work into customer and workload optimizations but there is lots more we can do to enable the broad ecosystem. I appreciate all the feedback and desire to engage with @AMD. We are committed to building a world-class open software stack. Lots planned for 2025. Happy holidays to all!"
They don’t fucking want to! Believing this is anything like a market is fucking religion.
They can spend their market cap by either:
1: issuing new shares worth their market cap, diluting existing shareholders to 50%.
2: Or borrow their market cap and pay interest by decreasing profits. "AMD operating margin for the quarter ending September 30, 2024 was 5.64%" so profits would be extremely impacted by interest repayments.
Either way your suggestion would be unlikely to be supported by shareholders.
> crush the next TSMC node on Apple levels
I would guess Apple is indirectly paying for the hardware (to avoid repatriating profits) or guaranteeing usage to get to the front of the line at TSMC. Good luck AMD competing with Apple: there's a reason AMD sold GlobalFoundries and there's a reason Intel is now struggling with their foundry costs.
And it comes across as condescending to assume you know better than a successful company.
What in God’s name do we pay these structured finance, bond-issue assholes 15% of GDP for if not to finance a sure thing like that?
It sure as hell ain’t for their taste in Charvet and Hermes ties, because the ones they pick look like shit.
Remember, most acquisitions fail. For the same reason, the likelihood of failure with your scenario seems high.
Do you really think nobody at AMD is aware of all the points made in this thread? That seems too bizarre to be true. There are probably some issues in upper management which could perhaps be fixed with some targeted hiring decisions, but do you really believe some random person on here would have a chance making that call?
There’s this meme that it can’t change on a dime and I believe that.
You could build this from scratch in a decade. JFK sent NASA to the moon in less time for comparable money.
If NVIDIA shareholders can’t come close? What fucking good are they? Why do our carrier battle groups guard their supply chain?
I say ship them Altman and an exaflop and watch their society corrupt itself at a fractal nature at machine speed.
Good fucking riddance. I see your fentanyl crisis: raise you Sam and a failure to ship GPT-5. Have fun with that.
“Cuda moat” is a misnomer. The PTX spec is relatively short (600 page pdf). Triton directly writes PTX, skipping cuda. Flash attention was created by a non nvidia employee without access to any of the secret sauce within Cuda or its libraries.
The hardware is just not as good, and no software can paper over its flaws.
Otherwise who will bet their firm / cash / career on new hardware without a successful track record.
Wrong conclusion. AMD is slower than NVidia, but not _that_ much slower. They are actually pretty cost-competitive.
The just need to do some improvements, and they'll be a very viable competitor.
And, if you’re training a model costing you millions, the last thing you need is a buggy, untested stack, breaking training or perhaps worse giving you noise that makes your models perform worse or increases training time.
By the time AMD gets usable out of the box at this point, NVidia will have moved further ahead.
Meanwhile, Nvidia hardware is expensive and still is in short supply. AMD might look quite tempting.
Meanwhile those libs release running CUDA on NVidia’s old and newest releases out of the box.
So no, it cannot be reused by others in production any more than my custom hacked car engine mod can be added by Ford to every car in existence.
Have you done any deep professional production work on any of these stacks? I have, and would never, ever put stuff like the stuff in the article in production. It’s no where near ready for production use.
> Getting reasonable training performance out of AMD MI300X is an NP-Hard problem.
2. AMD has always shipped a bad software stack, it is no different with AI.
https://x.com/HotAisle/status/1870984996171006035