Testing AMD's Giant MI300X
chipsandcheese.com
chipsandcheese.com
NVIDIA was there with 57 papers, a website dedicated to their research presented at the conference, a full day tutorial on accelerating deep learning, and ever present with shirts and backpacks in the corridors and at poster presentations.
AMD had a booth at the expo part, where they were raffling off some GPUs. I went up to them to ask what framework I should look into, when writing kernels (ideally from Python) for GPGPU. They referred me to the “technical guy”, who it turns out had a demo on inference on an LLM. Which he couldn’t show me, as the laptop with the APU had crashed and wouldn’t reboot. He didn’t know about writing kernels, but told me there was a compiler guy who might be able to help, but he wasn’t to be found at that moment, and I couldn’t find him when returning to the booth later.
I’m not at all happy with this situation. As long as AMDs investment into software and evangelism remains at ~$0, I don’t see how any hardware they put out will make a difference. And you’ll continue to hear people walking away from their booth, saying “oh when I win it I’m going to sell it to buy myself an NVIDIA GPU”.
Back then AMD/ATI were actually at the forefront on the GPGPU side - things like the early brook language and CTM lead pretty quickly into things like OpenCL. Lots of work went on using the xbox360 gpu in real games for GPGPU tasks.
But CUDA steadily improved iteratively, and AMD kinda just... stopped developing their equivalents? Considering a good part of that time they were near bankruptcy it might have not have been surprising though.
But saying Nvidia solely kicked off everything with CUDA is rather a-historical.
I've been writing CUDA since 2008 and it doesn't seem that different to me. They even still use some of the same graphics in the user guide.
I wasn't so much that they stopped developing, rather they kept throwing everything out and coming out with new and non backwards compatible replacements. I knew people working in the GPU Compute field back in those days who were trying to support both AMD/ATI and NVidia. While their CUDA code just worked from release to release and every new release of CUDA just got better and better, AMD kept coming up with new breaking APIs and forcing rewrite and rewrite until they just gave up and dropped AMD.
AMD's software investments have begun in earnest a few years ago, but AMD really did progress more than pretty much everyone else aside from NVidia IMO.
AMD further made a few bad decisions where they "split the bet", relying upon Microsoft and others to push software forward. (I did like C++ Amp for what its worth). The underpinnings of C++Amp led to Boltzmann which led to ROCm, which then needed to be ported away from C++Amp and into CUDA-like Hip.
So its a bit of a misstep there for sure. But its not like AMD has been dilly dallying. And for what its worth, I would have personally preferred C++ Amp (a C++11 standardized way to represent GPU functions as []-lambdas rather than CUDA-specific <<<extensions>>>). Obviously everyone else disagrees with me but there's some elegance to parallel_for_each([](param1, param2){magically a GPU function executing in parallel}), where the compiler figures out the details of how to get param1 and param2 from CPU RAM into GPU (or you use GPU-specific allocators to make param1/param2 in the GPU codespace already to bypass the automagic).
Last time I checked they have been trying to hire a ton of software engineers for improving the applied stacks (CV, ML, DSP, compute, etc) at the location near where I'm located.
It seems like there's a big push to improve the stacks but given that less than 10 years ago they were practically at death's door it's not terribly surprising that their software is in the state it is. It's been getting better gradually but quality software doesn't just show up over night and especially so when things are as complex and arcane as they are in the GPU world.
There is always financing, there are always people willing to go to the competitor at some wage, there is always a way if the leadership wants to.
If it was just a straight up fab bottleneck? Yeah maybe you buy that for a year or two.
“During Q1, Nvidia reported $5.6 billion in cost of goods sold (COGS). This resulted in a gross profit of $20.4 billion, or a margin profile of 78.4%.”
That’s called an “induced market failure”.
People love to pop-off on stuff they really know anything about. Let me ask you: what financing do you imagine is available? Like literally what financing do you propose for a publically traded company? Like do you realize they can't actually issue new shares without putting it to a shareholder vote? Should they issue bonds? No I know they should run an ICO!!!
And then what margins exactly? Do you know what the margin is on MI300? No. Do you know whether they're currently selling at a loss to win marketshare? No.
I would the happiest boy if hn, in addition to policing jokes and memes, could police arrogance.
Really? I must be reading a different language than English here
> There is always financing, there are always people willing to go to the competitor at some wage, there is always a way if the leadership wants to.
The issue was that they basically let everyone go who wasn't building hardware for their essential product lines (CPU & GPU) other than a skeleton crew to keep the software at least mostly functioning. And as much as this seems like it was a bad decision, AMD was probably weeks from bankruptcy by the time they got Zen out the door even despite doing this. Had they not done so, they'd almost certainly closed up entirely.
So for the last ~5 years minimum now they've been building back their software teams and trying to recuperate what they lost in institutional knowledge. That all takes time to do even if you hire back twice as many engineers as you lost.
And so now we are here. Things are clearly improving but nowhere near acceptable yet. But there's a trend of improvement.
How long am I supposed to wait, as my still-modern AMD GPU sits still-unsupported?
The anecdote above doesn't even sound like there's any improvement at all, let alone "clear" improvement.
And with Zen in 2017 and Zen+ in 2018 the counter is past six years at this point since the money gates opened wide.
Which GPU do you have? At least according to these docs, on linux the upper chunk of RDNA3 is supported officially but from experience, basically all 6xxx or 7xxxx cards are unofficially supported if you build it for your target arch. 5xxx cards get the short end of the stick and got skipped (they were a rough launch) but Radeon VII cards should also still be officially supported (with support shifting to unofficial status in the next release).
https://rocm.docs.amd.com/en/latest/compatibility/compatibil...
And given that ROCm is pretty core to AMD's support for the windows AI stack (via ONNX), you can assume any new GPUs released from here on out will be supported.
The unofficial support for so many cards is not a good situation either.
Edit: Actually, no, I know it's not that different, because some versions of ROCm largely work on RDNA1 if you trick them. They are just refusing to do the extra bit of work to patch over the differences.
https://www.reddit.com/r/ROCm/comments/1bd8vde/psa_rdna1_gfx...
I wish they had comprehensive support for basically all recent GPU releases but tbh I'd rather they focus on perfecting support for the current and upcoming generations than spread their efforts too thin.
And ideally with time backports to the older cards will come with time but it's really not a priority over issues on the current generation because those RDNA1 cards were never actually supported in the first place.
Financing is not the bottleneck. Organizational capacity might well be, though. As an organization, AMDs survival depended not on competing with nVidia but on competing with Intel. Now they are established, in what must be one of the greatest come from behind successes in tech history. 8 years ago, Intel was worth 80 times as much as AMD, today AMD has surpassed them:
https://www.financecharts.com/compare/AMD,INTC/summary/marke...
Stock isn't reality, but I wouldn't so easily assume that the team that led AMD to overtake Intel are idiots.
"Modular to bring NVIDIA Accelerated Computing to the MAX Platform"
https://www.modular.com/blog/modular-partners-with-nvidia-to...
Especially since CUDA is still rolling out new functionality and optimizations, so the goal posts will keep moving.
Corporations simply arn't interested in long term gains unless there's a straightforward path.
When you are a government agency, it’s more palatable to spend the budget in a way it results in employment of nationals and development of indigenous technologies.
Rational hyperscalers would just stop as soon as their tooling/workloads/models are functional on AMD hardware within an acceptable perf envelope - just like they already do with their custom silicon. Replicating CUDA is just unnecessary, expensive and time-consuming completionism; if some workloads require CUDA, they will be executed on Nvidia clusters that are part of the fleet.
It was this same game with x86 and ARM is eroding the former king’s place in the datacenter.
Another problem is simply that hiring (and keeping) top talent is really really hard. If you're smart enough to be a lead developer of AMDs core Machine Learning libraries, you can probably get hired at any number of other places, so why choose AMD.
I think the leadership gets it and understand the importance, I just don't think they (or really anybody) knows how to come up with a good plan to turn things around quickly. They're going to have to commit to at least a 5 year plan and lose money each of those 5 years, and I'm not sure they can or even want to fight that battle.
Absolutely. And when your mandate for this top talent is going to be "go and build something that basically copies what those other guys have already built", it is even harder to attract them, when they can go any place they like and work on something new.
> I think the leadership gets it and understand the importance, I just don't think they (or really anybody) knows how to come up with a good plan to turn things around quickly.
Yes, it always puzzles me when people think nobody at AMD actually sees the problem. Of course they see it. Turning a large company is incredibly hard. Leadership can give direction, but there is so much baked in momentum, power structures, existing projects and interests, that it is really tough to change things.
But for years I have heard the same things from so many people working in the field. "We hate Nvidia because they got it so right but are the only option."
Just like Intel, they have an outdated culture. IMHO they should start a software Skunk Works isolated from the company and have the software guys guide the hardware features. Not the other way around.
I wouldn't bet money on either of them doing this. Hopefully some other smaller, modern, and flexible companies can try it.
Their GPUs were re-designed to follow C++ memory model, and many NVidia engineers are seat at ISO C++, yet making CUDA the best way to run heterogenous C++. Something that Intel also realized, by acquiring CodePlay, key players in SYCL, and also employing ISO C++ contributors.
Then there are the Visual Studio and Eclipse plugins, and graphical debuggers that allow even to single step shaders if you so wish.
[0] https://tinygrad.org/ [1] https://github.com/tinygrad/tinygrad [2] https://x.com/realGeorgeHotz/status/1800932122569343043?t=Y6...
[0] https://github.com/tinygrad/tinygrad/issues/4301 [1] https://x.com/realAnthonix/status/1800993761696284676
For what it's worth, your post has cemented my decision to submit a few conference talks. I've felt too busy writing code to go out and speak, but I really should make time.
https://docs.cupy.dev/en/v13.2.0/install.html#using-cupy-on-...
This is the kind of stuff AMD keeps missing out, even OneAPI from Intel looks better in that regard.
It appears AMD initial strategy is courting the HPC crowd and hyperscalers, they have big budgets, lower support overhead and are willing and able to write code that papers-over AMDs not-great software while appreciating lower-than-Nvidia TCO. I think this this incremental strategy is sensible, considering where most of the money is.
As a first mover, Nvidia had to start from the bottom up; CUDA used to run only/mostly on consumer GPUs - AMD is going top-down, starting with high-margin DC hardware, before trickling down rack-level users, and eventually APUs later as revenue growth allows more re-investment.
They will fail if they go after the highest margin customers. Nvidia has every advantage and every motivation to keep those customers. They would need a trillion dollars in capital to have a chance imho.
It would be like trying to go after Intel in the early 2000s by trying to target server cpus, or going after the desktop operating system market in the 90s against Microsoft. Its aiming for your competition where they are strongest and you are weakest.
Their only chance to disrupt is to try to get some of the customers that Nvidia doesn’t care about, like consumer level inference / academic or hobbyist models. Intel failed when they got beaten in a market they didn’t care about, i.e mobile / small power devices.
All AMD would really need is for Nvidia innovation to stall. Which, with many of their engineers coasting on $10M annual compensation, seems not too far fetched
But I see no evidence that the strategy is wrong or failing. AMD is already powering a massive and rapidly growing share of Top 500 HPC:
https://www.top500.org/statistics/treemaps/
AMD compute growth isn't in places where people see it, and I think that gives a wrong impression. (Or it means people have missed the big shifts over the last two years.)
I use my university's "supercomputer" every now and then when I need lots of VRAM, and there are rarely many other users. E.g. I've never had to queue for a GPU even though I use only the top model, which probably should be the most utilized.
Also, I'd guess there can be nvidia cards in the grid even if "the computer" is AMD.
Of course it doesn't matter for AMD whether the compute is actually used or not as long as it's bought, but lots of theoretical AMD flops standing somewhere doesn't necessarily mean AMD is used much for compute.
That being said, there is a certain irony and schadenfreude in the AMD laptop being bricked from the thread root. The AMD engineers are at least aware that running a compute demo is an uncomfortable experience on their products. The consumer situation is not acceptable even if strategically AMD is doing OK.
In the cluster I'm using there's 36 nodes, of which 13 are currently not idling (doesn't mean they are computing). There are 8 V100 GPUs and 7 A100 GPUs and all are idling. Admittedly it's holiday season and 3AM here, but this it's similar other times too.
This is of course great for me, but I think the safer bet is that the typical load average of a "supercomputer" is under 0.10. And the less useful the hardware, the less will be its load.
I do see the phenomenon you describe on smaller university clusters, but these are not power users who know how to leverage HPC to the highest capacity. People in DOE spend their careers working to use as much as these machines as efficiently as possible.
Also this amount of GPUs is not sufficient for competitive pure ML research groups from what I have seen. The point of these small decentral underutilized resources is to have slack for experimentation. Want to explore ML application with a master student in your (non-ML) field? Go for it.
Edit: No idea how much of the total hpc market is in the many small instalks, vs the fewer large ones. My instinct is that funders prefer to fund large centralised infrastructure, and getting smaller decentralised stuff done is always a battle. But that's all based on very local experience, and I couldn't guess how well this generalises.
> It is a pretty safe bet that if someone builds a supercomputer there is a business case for it.
As I understand, most (95%+) of the market for supercomputers is gov't. If wrong, please correct. Else, what do you mean by "business case"?Despite the clichés, spending taxpayer money is really hard. In fact my impression is always that the fear that resources get misused is a major driver of the inefficient bureaucracies in government. If we were more tolerant of taxpayer money being wasted we could spend it more efficiently. But any individual instance of misuse can be weaponized by those who prefer for power to stay in the hands of the rich...
With the difficulty of spending taxpayer money, I fully agree. I even think HPC clusters are a bit of a symptom of this. It's often really hard to buy a beefy enough workstation of your own that would fit the bill, or to just buy time from cloud services. Instead you have to faff with a HPC cluster and its bureaucracy, because it doesn't mean extra spending. And especially not doing a tender, which is the epitome of the inefficiency caused by the paranoia of wasted spending.
I've worked for large businesses, and it's a lot easier to spend in those for all sorts of useless stuff, at least when the times are good. When the times get bad, the (pointless) bureaucracy and red tape gets easily worse than in gov organizations.
Because the users expect them to be renewed and improved. Otherwise the research can’t be done. None of our users tell us to buy new systems. But they cite us like mad, so we can buy systems every year.
The dynamics of this ecosystem is different.
I’m in that ecosystem. Access is limited, demand is huge. There’s literal queues and breakneck competition to get time slots. Same for CPU and GPU partitions.
They generally run at ~95% utilization. Even our small cluster runs at 98%.
Or let me ask you directly, can you name me one enterprise which would buy a super computer and wait 5+ years for it and fund the development of HW for it which doesn't exist yet? At the same time when the competition can deliver a super computer within the year with an existing product?
No sane CEO would have done Frontier or El Capitan. Such things work only with government funding where the government decides to wait and fund an alternative. But AMD is indeed a bit lucky that it happened or otherwise they wouldn't been forced to push the Instinct line.
In the commercial world, things work differently. There is always a TCO calculation. But one critical aspect since the 90s is SW. No matter how good the HW is, the opportunity costs in SW could force enterprises to use the inferior HW due to SW deployment. If vision computing SW in industry is supporting and optimized for CUDA or even runs only with CUDA then any competition has a very hard time penetrating that market. They first have to invest a lot of money to make their products equally appealing.
AMD makes a huge mistake and is by far not paranoid enough to see it. For 2 decades, AMD and Intel have been in a nice spot with PC and HPC computing requiring x86. It basically to this date has guaranteed a steady demand. But in that timeframe mobile computing has been lost to ARM. ML/AI doesn't require x86 as Nvidia demonstrates by combining their ARM CPUs into the mix but also ARM themselves want more and more of the PC and HPC computing cake. And MS is eager to help with OS for ARM solutions.
What that means is that if some day x86 isn't as dominant anymore and ARM becomes equally good then AMD/Intel will suddenly have more competition in CPUs and might even offer non-x86 solutions as well. Their position will therefore drop into yet another commodity CPU offering.
In the AI accelerator space we will witness something similiar. Nvidia has created a platform and earns tons of money with it by combining and optimizing SW+HW. Big Tech is great at SW but not yeat at HW. So the only logical thing to do is getting better at HW. All large Tech companies are working on their own accelerators and they will build their platform around it to compete with Nvidia and locking in customers all the same way. The primary losers in all of this will be HW only vendors without a platform, hoping that Big Tech will support them on their platforms. Amazon and Google have already shown today that they have no intention to support anything besides their platform and Nvidia (which they only must due to customer demand).
Our first deployment is 3x larger flops than Cheyenne and a fraction of the cost.
They're able to get the high end customers, and this strategy works because they can sell the high end customers high end parts in volume without having to have a good software stack; at the high end, the customers are willing to put in the effort to make their code work on hardware that is better in dollars/watts/availability or whatever it is that's giving AMD inroads into the supercomputing market. They can't sell low end customers on GPU compute without having a stack that works, and somebody who has a small GPU compute workload may not be willing or able to adapt their software to make it work on an AMD card even if the AMD card would be a better choice if they could make it work.
Every framework, library, demo, tool, and app is going to use CUDA forever and ever while some “account manager” at AMD takes a government procurement officer to lunch to sell one more supercomputer that year.
Their quarterly data centre revenue is now $22.6B! Even assuming that it immediately levels off, that's $90B over the next 12 months.
If it merely doubles, then they'll hit a total of $1T in revenue in about 6 years.
I'm an AI pessimist. The current crop of generative LLMs are cute, but not a direct replacement for humans in all but a few menial tasks.
However, there's a very wide range of algorithmic improvements available, which wouldn't have been explored three years ago. Nobody had the funding, motivation, or hardware. Suddenly, everyone believes that it is possible, and everyone is throwing money at the problem. Even if the fruits of all of this investment is just a ~10% improvement in business productivity, that's easily worth $1T to the world economy over the next decade or so.
AMD is absolutely leaving trillions of dollars on the table because they're too comfortable selling one supercomputer at a time to government customers.
Those customers will stop buying their kit very soon, because all of the useful software is being written for CUDA only.
Who can even afford to buy that much product? Are you expecting Apple, Microsoft, Alphabet, Amazon, etc to all dump 100% of their cash on Nvidia GPUs? Even then that doesn't get you to a trillion dollars
This kind of AI capital investment seems to have helped them improve the feed recommendations, doubling their market cap over the last few years. In other words, they got their money back many times over! Chances are that they're going to invest this capital into B100 GPUs next year.
Apple is about to revamp Siri with generative AI for hundreds of millions of their customers. I don't know how many GPUs that'll require, but I assume... many.
There's a gold rush, and NVIDIA is the only shovel manufacturer in the world right now.
Right, which means you need about a trillion dollars more to get to a trillion dollars. There's not another 100 Metas floating around.
> Apple is about to revamp Siri with generative AI for hundreds of millions of their customers. I don't know how many GPUs that'll require, but I assume... many.
Apple also said they were doing it with their silicon. Apple in particular is all but guaranteed to refuse to buy from Nvidia even.
> There's a gold rush, and NVIDIA is the only shovel manufacturer in the world right now.
lol no they aren't. This is literally a post about AMD's AI product even. But Apple and Google both have in-house chips as well.
Nvidia is the big general party player, for sure, but they aren't the only. And more to the point, exponential growth of the already largest player for 6 years is still fucking absurd.
If AI can improve productivity by just 1% then that is $2T more. If it costs $1T in NVIDIA hardware then this is well worth it.
That is assuming Nvidia can capture the value and doesn't get crushed by commodity economics. Which I can see happening and I can also see not happening. Their margins are going to be under tremendous pressure. Plus I doubt Meta are going to be cycling all their GPUs quarterly, there is likely to be a rush then settling of capital expenses.
All evidence is that “more is better”. Everyone involved professionally is of the mind that scaling up is the key.
However, like you said, just a single invention could cause the AI winds to blow the other way and instantly crash NVIDIA’s stock price.
Something I’ve been thinking about is that the current systems rely on global communications which requires expensive networking and high bandwidth memory. What if someone invents an algorithm that can be trained on a “Beowulf cluster” of nodes with low communication requirements?
For example the human brain uses local connectivity between neurons. There is no global update during “training”. If someone could emulate that in code, NVIDIA would be in trouble.
CUDA relevance in the industry is so big now, that NVidia has several WG21 seats, and helps driving heterogenous programming roadmap for C++.
https://www.cgchannel.com/2023/11/otoy-releases-first-public...
https://www.cgchannel.com/2023/11/otoy-unveils-the-octane-20...
And then again, there are plenty of other cases where Pytorch makes absolute no sense in GPU, which was the whole starting point.
I said that if they wanted to support AMD they would use the closest-to-metal API possible, and your links prove that this is exactly their mindset - preferring a lower level more performant API to a higher level more portable one.
For many people the tradeoffs are different and ability to write code quickly and iterate on design makes more sense.
I don't see how the same can work here. HIP isn't it right now (every time I try, anyway).
They are already powering the most powerful supercomputers, so I guess you’re right.
Oh, by coincidence, the academic crowd is the primary user of these supercomputers.
Pure luck.
I don't agree with this at all! Give me something that I can easily prototype at home and then quickly scale up at work!
That was a marketing guy BTW. I don't think they realize their marketing strategies suck.
We are thrilled to announce that Hot Aisle Inc. proudly volunteered our system for Chips and Cheese to use in their benchmarking and performance showcase. This collaboration has demonstrated the exceptional capabilities of our hardware and further highlighted our commitment to cutting-edge technology.
Stay tuned for more exciting updates!
Funny that HN doesn't like my comment for some reason though.
Previously, this stuff was only available to HPC applications. We're trying to get these into the hands of more developers. Our view is that this is a great way to foster the ecosystem.
Our simple and competitive pricing reflects this as well.
Personally, I couldn't care less about the quality of copy. I do care about having access to similar hardware in the future.
Much like with AI, Nvidia has the software side of GPU production rendering locked down tight though so that's just as much of an uphill battle for AMD.
They currently only use the GPU mode for quick iteration on relatively small slices of data though, and then switch back to CPU mode for the big renders.
In the case of Stadia, however, failing to develop this was like a sports team not playing any home games. One way of thinking about the current crisis of the games industry and VR is that building 3-d worlds is too expensive and a major part of it is all the shoehorning tricks the industry depends on. Better hardware for games could be about lowering development cost as opposed to making fancier graphics but that tends to be a non-starter with companies whose core competence is getting 1000 highly-paid developers to struggle with difficult to use tools and the idea you could do the same with 10 ordinary developers is threatening to them.
I am thinking of an entire datacenter purpose-built to host a single game world, with edge locations handling the last mile of client-side prediction, viewport rendering, streaming and batching of input events.
We already have a lot of the conceptual architecture figured out in places like the NYSE and CBOE - Processing hundreds of millions of events in less than a second on a single CPU core against one synchronous view of some world. We can do this with insane reliability and precision day after day. Many of the technology requirements that emerge from the single instance WoW path approximate what we have already accomplished in other domains.
I have been thinking about if the the compute could go right in cellphone towers but this would take it up a notch.
Ironically the main type that'd still exist would be the vision-based external AI-powered target-highlighting and aim/fire assist.
The display is analysed and overlaid with helpful info (like enemies highlighted) and/or inputs are assisted (snap to visible enemies, and/or automatically pull trigger.)
I'm betting with current hardware and some clever tricks, we can resolve full production frames in real-time rates.
But like many have said considering AMD was almost bankrupt their performance is impressive. This really speaks for their hardware division. If only they could get the software side of things fixed!
Also I wonder if NVIDIA has an employee of the decade plaque for CUDA. Because CUDA is the best thing that could’ve happened to them.
We can't possibly hope to run the kinds of models that run on 192GB of VRAM at home.
So maybe Apple is happy to sell huge GPUs like that but the government will probably put it under export controls like A100 and H100 already is
How many companies use Macs for ML work instead of Nvidia and Cuda?
Besides, you'd be well served with a Mac as a development desktop anyway.
At least until someone makes an MI300A workstation.
What you alone do at home, is irelevant for the ML market as a whole, along with your Mac Mini, as you alone won't move the market, and the companies serious about ML are all-in on Nvidia and CUDA compatible code for mass deployment.
I can also get to run some NNs on some microcontroler, but my hoppy project won't move the market, and that's what I was talking about, the greater market, not your hobby project.
Side note: as someone who has been into machine learning for over 10 years, let me tell ya us hobbyists and researchers hunger for compute and memory.
VRAM isn't everything.....I am well aware but certain workflows really do benefit from heaps of vram like vfx and cad and CFD. I realize that the dream of upgradable GPUs where I can upgrade the different components just like you do on the computer. Computer is slow, then upgrade ram or storage or get a faster chip that uses the same socket. GPU could possibility see modularity with the processor the vram etc.
Level1Tech has some great videos about how PCIe is the future...where we can connect systems together using raw PCI lanes, which is similar to how nvidia Blackwell servers communicate to other servers in the rack.
I'm looking to build a mini-ITX system with 256GB of RAM for my next build. DDR5 spec can support that in 2 modules, but nobody makes them yet. No need for a GPU, I'm looking to the AMD APUs which are getting into the 50TOPs range. But yes, RAM seems to be the limiting factor. I'm a little surprised the memory companies aren't pushing harder for consumers to have that capacity.
It is of course RDIMM, but you didn't specify what memory type you were looking at.
Or there's openmp or hip. In extremis opencl.
I think the language stack is fine at this point. The moat isn't in cuda the tech. It's in code running reliably on nvidia's stack, without things like stray pointers needing a machine reboot. Hard to know how far off robust rocm is at this point.
ROCm can be spotty, especially on consumer cards, but for many models it does seem to work on their more expensive models. It may be worth it spending a few hours/days/weeks to work around the peculiarities of ROCm given the cost difference between AMD and Nvidia in this market segment.
This all stands or falls with how well AMD can get ROCm to work. As this article states, it's nowhere near ready yet, but one or two updates can turn AMD's accelerators from "maybe in 5-10 years" to "we must consider this next time we order hardware".
I also wonder if AMD is going to put any effort into ROCm (or a similar framework) as a response to Qualcomm and other ARM manufacturers creaming them on AI stuff. If these Copilot PCs take off, we may see AMD invest into their AI compatibility libraries because of interest from both sides.
That's mostly on Microsoft's DirectML though. I'm not sure whether AMD's implementation is based on ROCm (doubt it).
I would have thought AMD could have scrambled to fix their bugs, at least the matmul related ones, scrambled to shore up torch compatibility or whatever was needed for LLM training, and pushed something out the door that might not have been top-of-market but could at least have taken advantage of the opportunity provided by 80% margins from team green. I thought the green moat was maybe a year wide and tens of millions deep (enough for a team to test the bugs, a team to fix the bugs, time to ramp, and time to make it happen). But here we are, multiple years and trillions in market cap delta later, and AMD still seems to be completely non-viable. What happened? Did they go into denial about the bugs? Did they fix the bugs but the industry still doesn't trust them?
Other people think it's buggy and useless because that's the experience on some other platforms.
This state of affairs isn't great. It could be worse but it could certainly be much better.
"One of the things that you mentioned earlier on software, very, very clear on how do we make that transition super easy for developers, and one of the great things about our acquisition of Xilinx is we acquired a phenomenal team of 5,000 people that included a tremendous software talent that is right now working on making AMD AI as easy to use as possible."
Xilinx dev tools are awful. They are the ones who had Windows XP as the only supported dev environment for a product with guaranteed shipments through 2030. I saw Xilinx defend this state of affairs for over a decade. My entire FPGA-programming career was born, lived, and died, long after XP became irrelevant but before Xilinx moved past it, although I think they finally gave in some time around 2022. Still, Windows XP through 2030, and if you think that's bad wait until you hear about the actual software. These are not role models of dev experience.
In my, err, uncle? post I said that I was confused about where AMD was in the AI arms race. Now I know. They really are just this dysfunctional. Yikes.
> He tries to get a comment on the (in hindsight) not great design tradeoffs made by the Cell processor, which was hard to program for and so held back the PS3 at critical points in its lifecycle. It was a long time ago so there's been plenty of time to reflect on it, yet her only thought is "Perhaps one could say, if you look in hindsight, programmability is so important". That's it! In hindsight, programmability of your CPU is important! Then she immediately returns to hardware again, and saying how proud she was of the leaps in hardware made over the PS generations.
> He asks her if she'd stayed at IBM and taken over there, would she have avoided Gerstner's mistake of ignoring the cloud? Her answer is "I don’t know that I would’ve been on that path. I was a semiconductor person, I am a semiconductor person." - again, she seems to just reject on principle the idea that she would think about software, networking or systems architecture because she defines herself as an electronics person.
> Later Thompson tries harder to ram the point home, asking her "Where is the software piece of this? You can’t just be a hardware cowboy ... What is the reticence to software at AMD and how have you worked to change that?" and she just point-blank denies AMD has ever had a problem with software. Later she claims everything works out of the box with AMD and seems to imply that ROCm hardly matters because everyone is just programming against PyTorch anyway!
> The final blow comes when he asks her about ChatGPT. A pivotal moment that catapults her competitor to absolute dominance, apparently catching AMD unaware. Thompson asks her what her response was. Was she surprised? Maybe she realized this was an all hands to deck moment? What did NVIDIA do right that you missed? Answer: no, we always knew and have always been good at AI. NVIDIA did nothing different to us.
> The whole interview is just astonishing. Put under pressure to reflect on her market position, again and again Su retreats to outright denial and management waffle about "product arcs". It seems to be her go-to safe space. It's certainly possible she just decided to play it all as low key as possible and not say anything interesting to protect the share price, but if I was an analyst looking for signs of a quick turnaround in strategy there's no sign of that here.
not expecting a heartfelt postmortem about how things got to be this bad, but you can very easily make this question go away too, simply by acknowledging that it's a focus and you're working on driving change and blah blah. you really don't have to worry about crushing some analyst's mindshare on AMD's software stack because nobody is crazy enough to think that AMD's software isn't horrendously behind at the present moment.
and frankly that's literally how she's governed as far as software too. ROCm is barely a concern. Support base/install base, obviously not a concern. DLSS competitiveness, obviously not a concern. Conventional gaming devrel: obviously not a concern. She wants to ship the hardware and be done with it, but that's not how products are built and released in 2020 anymore.
NVIDIA is out here building integrated systems that you build your code on and away you go. They run NVIDIA-written CUDA libraries, NVIDIA drivers, on NVIDIA-built networks and stacks. AMD can't run the sample packages in ROCm stably (as geohot discovered) on a supported configuration of hardware/software, even after hours of debugging just to get it that far. AMD doesn't even think drivers/runtime is a thing they should have to write, let alone a software library for the ecosystem.
"just a small family company (bigger than NVIDIA, until very recently) who can't possibly afford to hire developers for all the verticals they want to be in". But like, they spent $50b on a single acquisition, they spent $12b in stock buybacks over 2 years, they have money, just not for this.
Heck I think it is being used to run ChatGPT 3.5 and 4 services.
Not sure why they haven’t managed to execute on that yet, but the partners must be pretty motivated now, right? I’m sure they don’t love doing business at Nvidia’s leisure.
cheaper is true, but less power hungry is absolutely not true, which is kind of my point.
In any case they're only slightly behind, not crazy far behind like Intel is.
When you're providing fab designs at that scale, it makes a lot more sense to folks that companies would be willing to try a more affordable option to nVidia hardware.
My bet is that AMD figures out a service-able solution for some (not all) workloads that isn't ground breaking, but affordable to the clients that want an alternative. That's usually how this goes for AMD in my experience.
Nobody forget that, just that those console chips are super low margins, which is why Intel and Nvidia stopped catering to that market after the Xbox/PS3 generations and only AMD took it up because they were broke and every penny mattered to them.
Nvidia did a brief stint with the Shield/Switch because they were trying to get into the Android/ARM space and also kinda gave up due to the margins.
Among the gamer community the discussion of this being the last generation keeps poping up.
[0] - Nintendo is more than happy to keep redoing their hit franchaises, in good enough hardware.
I'm curious if that's just because they can't get enough Nvidia GPUs or if the price/performance is actually that much better.
Think of it this way: AMD is pretty good at hardware, so there's no reason to think that the raw difference in terms of flops is significant in either direction. It may go in AMD's favor sometimes and Nvidia's other times.
What AMD traditionally couldn't do was software, so those AMD GPUs are sold at a discount (compared to Nvidia), giving you better price/performance if you can use them.
Surely Microsoft is operating GPUs at large enough scale that they can pay a few people to paper over the software deficiencies so that they can use the AMD GPUs and still end up ahead in terms of overall price/performance.
For example, for bitandbytes (a common dependency in LLM world) there's a ROCm fork that the AMD maintainers are trying to merge in (https://github.com/TimDettmers/bitsandbytes/issues/107). Meanwhile an Intel employee merged a change that made a common device abstraction (presumably usable by AMD + Apple + Intel etc.).
There's a lot of that right now - super popular package that is CUDA-only is navigating how to make it work correctly with any other accelerator. We just need more information on what is supported.
Has this returned? Because for dual gpu/cpu workloads (alpha zero, etc) that would deliver effective “infinite bandwidth” between gpu and cpu. Using an apu of course gets you huge amounts of slowish memory. But being some to fling things around with abandon would be an advantage, particularly for development.
On the GPU there are additional pointer types for different local memory, e.g. LDS is a uint16_t indexing from zero. But even there you can still have a single pointer to "somewhere" and when you store to it with a single flat addressing instruction the hardware sorts out whether it's pointing to somewhere in GPU stack or somewhere on the CPU.
This works really well for tables of data. It's a bit of a nuisance for code as the function pointer is aimed at somewhere in memory and whether that's to some x86 or to some gcn depends on where you got the pointer from, and jumping to gcn code from within x86 means exactly what it sounds like.
They had some fancy marketing name for it at the time. But it wasn't on all chips, it should have been. Even if it was dog slow between PCIe GPU and CPU the unified interface would have been the right way to go. Also, amenable to automated scheduling.
The point still stand though, I want entirely unified GPU and CPU memory.
Fundamentally if you've got separate blocks of memory tied together by pcie then it's either annoying copying data across or a potential performance problem doing it behind the scenes.
A single block of memory that everything has direct access to is much better. It works very neatly on the APU systems.
Well, as I said that's amenable to automated planning.
But what I really, really want is a nice APU with 512GB+ of memory that both the CPU and GPU can access willy nilly.
The MI300A is an APU with 128gb on the package. They come in four socket systems, that's 512gb of cache coherent machine with 96 fast x64 cores and many GCN cores. Quite like a node from El Capitan.
I'm delighted with the hardware and not very impressed with the GPU offloading languages for programming it. The GCN and x64 cores are very much equal peers on the machine, the asymmetry baked into the languages grates on me.
(on non-apu systems, moving the data around in the background such that the latency is hidden is a nice idea and horrendously difficult to do for arbitrary workloads)
> Even if it was dog slow between PCIe GPU and CPU the unified interface would have been the right way to go
That is actually what happened. You can directly access pinned cpu memory over pcie on discrete gpus.
> Taking LLaMA 3 70B as an example, in float16 the weights are approximately 140GB, and the generation context adds another ~2GB. MI300X’s theoretical maximum is 5.3TB/second, which gives us a hard upper limit of (5300 / 142) = ~37.2 tokens per second.
Each weight is a FP16 float which is 2 Bytes worth of data, you have 70B tokens, so the total amount of data the weights take up is 140GB then you have a couple extra GBs for the context.
Then to figure out the theoretical tokens per second you just divide the amount of memory bandwidth, 5300GB/s in MI300X's case, by the amount of data that the tokens and context take up so 5300/142 which is about 37 tokens per second.
In which case, if the model weights are fitting in the VRAM and are already loaded, why does the bandwidth impact the rate of tok/s?
The weights are not really reused. Which means they are never in registers, or in L1/L2/L3 caches. They are always in VRAM and always need to be loaded back in again.
However, if you are batching multiple separate inputs you can reuse each weight on ech input, in which case you may not be entirely bandwidth bound and this analysis breaks down a bit. Basically you can't produce a single stream of tokens any faster than this rate, but you can produce more than one stream of tokens at this rate.
I think they mean 37.2 forward passes per second. And at 4008 tokens per second (from "LLaMA3-70B Inference" chart) it means they were using a batch size of ~138 (if using that math, but probably not correct). Right?
The MI300X does memory bandwidth better than anything else by a ridiculous margin, up and down the cache hierarchy.
It did not score very well on global atomics.
So yeah, that seems about right. If you manage to light up the hardware, lots and lots of number crunching for you.
That means e.g. 8xH100 with TensorRT-LLM / vLLM vs 8xMI300X with vLLM running many concurrent requests with reasonable # of input and output tokens. Ran both in fp8 and fp16.
Most of the benchmarks I've seen had setups that no one would use in production. For example running on a single MI300X or 2xH100 -- this will likely be memory bound, you need to go to higher batch sizes (more VRAM) to be compute bound to properly utilize these. Or benchmarking requests with unrealistically low # of input tokens.
Most kernels are super quick to rewrite, and higher level abstractions like PyTorch and JAX make dealing with CUDA a pretty rare experience for most people making use of large clusters and small installs. And if you have the money to build a big cluster, you can probably also hire the engineers to port your framework to the right AMD library.
The world has changed a lot!
The bigger challenge is that if you are starting up, why in the world would you give yourself the additional challenge of going off the beaten path? Its not just CUDA but the whole infrastructure of clusters and networking that really gives NVIDIA an edge, in addition to knowing that they are going to stick around in the market, whereas AMD might leave it tomorrow.
Edit: Or HIPIFY, a tool made by AMD as a translation for CUDA. https://github.com/ROCm/HIPIFY/blob/amd-staging/README.md
"When it is all said and done, MI300X is a very impressive piece of hardware. However, the software side suffers from a chicken-and-egg dilemma. Developers are hesitant to invest in a platform with limited adoption, but the platform also depends on their support. Hopefully the software side of the equation gets into ship shape. Should that happen, AMD would be a serious competitor to NVIDIA."
AMD can offer this product at a 40% discount and still make money tells you all you need to know.
History has shown that this idea is not as crazy as it sounds.