AMD Expands AI Product Lineup with GPU-Only Instinct Mi300X with 192GB Memory
anandtech.com
anandtech.com
At the same time, the cost is outrageous. Its not a low volume product.
Also, a 48GB 4090 (or 3090) would be trivial. So would a 48GB 7900. Its not done for purely anticompetitive reasons, one that AMD and Nvidia are unfortunately happy to go along with.
48Gb 3090: https://www.nvidia.com/en-us/design-visualization/rtx-a6000/
https://www.bhphotovideo.com/c/product/1753962-REG/pny_vcnrt...
As most economics 101 lecture would tell you, price is determined by supply and demand and in this case Nvidia essentially maximizing their profit based on the demand and segmentation of the market.
Is this why Intel started taking dGPU production more seriously in recent years?
Nvidia and AMD used to "let" their manufacturers double up the vram on gaming cards, but no more. Its pure, collusive, anticompetitive market segmentation.
https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuec...
One might speculate he’s perhaps pivoting to Intel. They’re not well-developed in application terms but that’s a piece he can develop, and actually with OneAPI that’s a lot of potential bang for the buck. And intel has actual ML accelerators and has some relatively powerful GPGPU stuff and is currently in a position of being forced to offer a lot of bang for buck to drive adoption, all of which makes sense for what he’s trying to do.
But AMD wants you to basically write and debug their runtime for them and nah not worth it, after fighting the installer on the official system config and then filing a couple bugs for the demo apps reproducibly crashing the kernel it’s just not worth the time.
ROCm is unserious even when you’re operating on supported hardware. This is the experience most people have with it.
You know, it doesn't surprise me that's what they're doing, because the HPC people aren't paying megabucks and then tolerating something that crashes the kernel... but people's problems/experiences aren't invalid either, and the delta is they're not running the same software. I don't like it but it makes sense.
1. they're a publicly listed company operating in a free market in an industry they themselves helped develop who's purpose is to make returns for their investors, not be liked by the public, they don't need to justify their pricing to their buyers. They're not selling essentials for survival like insulin, baby formula, or housing, they can charge as much as the market will bear for their consumer electronics products. Don't like the pricing? Don't buy it. Simple. Buy from the competition instead or older generations off the second hand market that fit your budget.
2. the price justification is that cutting edge silicon is and will always be in short supply, and buyers of the silicon in A100 form, like datacenters, use it to make money, therefore it's an investment that will yield returns, therefore they can justify spending way more money to outbid the gamers who buy the same silicon in RTX 4090 form and don't use it to make money but use it to play games therefore for them the product is worth less and it makes Nvidia smaller margins than what selling it to datacenters can. It's basic price segmentation that's been going on for decades.
I also don't like the GPU pricing situation but that's the market reality I can't change and downvoting the messenger won't change it either. My 2 cents.
AMD has been consistently making bad business decisions in the GPU space for years now. What's new? They're pricing their hardware near NVidia but with less performance and way less features and poor track record, especially for the ML users.
Thinking that Nvidia is some evil villain doing it to consumer out of spite, when they're just doing what any other company in their dominant position would do: charge as much as the market will bear.
Apple also doesn't have to charge you $200 for configuring the 512GB SSD over the 256 SSD which only costs them an extra $5 NAND chip, but they do it because they can. So does Nvidia and any other company in a dominant position with virtually no competition.
You might be reading into their comment too much, they don't really blame NVidia, just wish for more competition. IMHO you're being downvoted because you come across as needlessly confrontational to an innocent comment.
Maybe I'm wrong and OP really didn't understand the curent supply/demand issues which is why I gave a lengthy explanation of how it looks.
LLms running on standalone affordable boxes could become a ubiquitous home accessory, not to mention the potential for things like being able to leverage learning from widespread use to retrain the models, like Tesla does with their fleet.
Yes, exactly like all the big tech companies with vast departments for market research that hired a shitload of people, then when economy turned, they started doing layoffs. A.k.a short term profit seeking.
It's a dick move but it's not illegal (at least not now). Their purpose is to make greater returns for their investors, not make affordable ML hardware for every consumer.
>LLms running on standalone affordable boxes could become a ubiquitous home accessory
They could be, but why is it Nvidia's fault if they don't happen? They're in the business of making money, not fulfilling dreams, and currently they can barely fulfill their orders for the datacenter customers. Consumers could move the LLM stuff on cheaper Apple HW or Intel or AMD SoCs if they can't outbid the datacenter companies for Nvidia silicone.
Seems like a market opening for Nvidia's competitors to price cut them. If they don't exploit it and let Nvidia dominate, it's their fault. It's not Nvidia's fault their competitors were asleep at the wheel since they launched CUDA in 2007 and were helping researchers put their GPUs to use for parallel computing for over 15 years.
Something like chrome : chromium dynamic would work wonders for bringing a sort of standard, not to mention more minds could create better and better implementations.
By that time nvidia already had an advantage with CUDA and their opencl support always lagged.
I mean, sure, you could compile your stuff to XLA or just, I dunno, set up 800 of these cards and train the whole thing on ROCM. But then would you really, really use AMD instead of some TPUs?
AMD made Machine Learning either impossible, unsupported or a chore on most of their hardware for years. Their stack sucks and no one wanted to implement it, and in fact they rarely supported most of their own GPUs.
Yes you have GPUs. But we need also drivers. And software. People have been yelling at AMD about this for years and years.
Instead, AMD has made it clear that Machine Learning is not a priority for the company. Hence, they deserve the lack of traction. Investing in AMD hardware for ML has literally been a mistake at every point in recent history. Imagine if you bought a bunch of (insert last Gen card which is no longer supported by their stack) how dumb you'd look.
Releasing a GPGPU card now? Honestly, why bother? No one is gonna buy it.
NVIDIA is about to walk off with a trillion dollars because nobody at AMD “gets it”.
With no meaningful competition, NVIDIA will gouge as hard as they can. Such as charging $50K for a card that’s not too different to a 4090 but with more memory.
I think they "get it" OK. Whether or not they can formulate a viable strategy and execute it is one question, but they get the idea that "AI is important" and they know where they stand vis-a-vis NVIDIA.
This is another reason I'm willing to invest some time and money in working with AMD products for AI/ML. History has shown us their ability to go toe-to-toe with a seemingly unassailable industry titan before, and they came out in pretty good shape then.
https://www.forbes.com/sites/iainmartin/2023/05/31/lisa-su-s...
ROCm is getting there, slowly.
So basically, I'm betting that AMD has had (or is having) a change of heart and is genuinely committed to AI/ML on their products. Time may ultimately prove me wrong, but so be it if that proves to be the case.
And FWIW, one reason for my optimism is that, whatever you think about the state of ROCm today, they are clearly investing heavily into the platform and continually working on it. You can see that just from looking at the activity on their Github repos:
https://github.com/orgs/ROCmSoftwarePlatform/repositories
There is constant activity and has been for some time, which I take as a good sign. Yes, it's just one signal among many that one could consider, but I think it's an important one.
My example about LLMs was just to show that AMD is simply not part of the conversation. Three month before you could have made the same point about another approach.
And still, if you'd actually have to risk money, you and me both know you'd never invest in AMD hardware for AI or start developing on it something high stakes.
I mean, look at geohot he tried and just gave up entirely and AMD.
I assume you meant "you" in the royal sense. Since I personally am, in point of fact, investing in AMD hardware for AI. Yeah, it's a gamble, but that's my style. And I have to admit, part of it is driven by ideological reasons (ROCm being open source) and a simple desire to support AMD since I want them to become a serious competitor to NVIDIA. I believe that's an outcome that would benefit everybody.
I mean, look at geohot he tried and just gave up entirely and AMD.
I have to admit, I don't find that particularly compelling in any regard. Nothing against geohot, he's clearly a smart dude, but... I'm not judging a company based on his interactions with them. shrug
So for now, I don't see AMD getting any traction. And apparently, the quality of the drivers has not been improving. Time will tell.
Beyond that, it's just going to be wherever the research takes me. And not everything is necessarily going to be particularly suited for running on ROCm. I'll use CUDA on NVIDIA hardware as and when needed, and I'm also open to dabbling in low-level hardware stuff and playing around with DSP's, analog computing, FPGA's, etc. One thing I want to work with a bit is some Spiking Neural Network stuff, neuromorphic approaches, etc. So not sure where this will end up.
Heck, it's possible that some of the stuff I work with will turn out to be well suited for running on a plain old CPU, which is one reason I spec's a relatively high-end Ryzen CPU and 64GB of system RAM for these machines, which otherwise wouldn't necessarily be all that important for "pure" GPU computation.
Maybe I'll even wind up going "old school" and building myself a Beowulf cluster! So I can finally answer the age old question about "imagine a Beowulf cluster of these..."
There's a language called HIP which is a fairly close approximation to cuda. You can probably convert one to the other with determination and regex. The GPUs themselves are fundamentally different in ways that hopefully don't matter to your application (warp synchronisation is the big one in my opinion, but I suspect cuda applications ignore it and just live with the race conditions).
Cuda is great, but it's not strictly necessary for much of the latest AI / ML developments.
It may be possible to use it with consumer GPUs anyway, but many won't try because it's not officially supported.
https://rocm.docs.amd.com/en/latest/release/gpu_os_support.h... https://developer.nvidia.com/cuda-gpus
Something like the way chrome vs chromium is, or even a foundation like the linux foundation, where you have multiple distros contributing packages/etc back into the ecosystem.
I think cloud providers love exclusivity(Nvidia MSRP is significantly higher than it is available to clouds) and based on pricing compared to competitors like lambdalabs they have highest profit margin on GPU instances. Also based on availability, they likely have the highest utilisation. They definitely wouldn't want to commoditize the space. Google already has TPU that they could scale and sell to everyone but it would make the margins significantly smaller if they do it.
AMD GPU acceleration via CLBlast was merged back in mid-May in llama.cpp:master - it works and gives a boost (although not all AMD GPUs have been tuned for CLBLast - this is something that AMD should be doing tbt: https://github.com/RadeonOpenCompute/ROCm/issues/2161)
There is also a hipBLAS fork, which is slightly faster (~10% on my old Radeon VII) which maybe someone at AMD should be supporting to make its way into master: https://github.com/ggerganov/llama.cpp/pull/1087
I'll also note that exllama merged ROCm support and it runs pretty impressively - it runs 2X faster than the hipBLAS llama.cpp, and in fact, on exllama, my old Radeon VII manages to run inference >50% faster than my old (roughly equal class/FP32 perf) GTX 1080 Ti (GCN5 has 2X fp16, and 2X memory bandwidth, so there's probably more headroom even) for a relatively easy port. That's really impressive: https://github.com/turboderp/exllama/pull/7
(It's worth noting that for the latter, all you need to do is install the ROCm version of PyTorch and "it just works," which is refreshing: https://pytorch.org/get-started/locally/)
More like a dozen of engineers over a couple of weeks. Make sure popular LLMs can run on their hardware, as well as Stable Diffusion and other popular projects, and then they will see consumers flock to their hardware.
With consumers I mean something like what "gamers" used to be during all these past decades (and still are), those who won't be using it for business cases (what Quadro used to be) but for their hobby; this, but oriented towards AI, which currently is limited to LLM, image generation and ASR.
If they focus on this, the community will start helping, maybe even cleaning up their mess of repositories they have on GitHub.
There are two benefits over Nvidia: They have more VRAM and they have an open source software stack.
They just need to get the basics working for all those hobbyists, those who want to run little projects on their home hardware.
As for a Hackintosh, I'd imagine an Nvidia Grace or Grace Hopper as a good option, even though the 500+W TDP would require a MacPro-sized heatsink. And it can have up to 960GB of RAM.
edit: got a lot of details confused between the MI300 and the Grace.
These things are so interesting it's a shame they aren't cheap.
https://videocardz.com/newz/intel-arrow-lake-p-with-320eu-gp...
(Sorry, I cannot find the AMD rumor link atm)
But TBH the hybrid design is less interesting than you think, just because nothing really takes advantage of it. Hence Intel canceled their datacenter APU in favor of a pure Falcon Shores GPU due to a lack of interest from customers.
Looks like someone ported llama to apples metal v3 already and are getting 5 tok/s on a 65b model.
The tape out cost would be huge, the die would be huge. Either the mobo/socket would be super expensive and niche, or consumers would be pissed about non expandable RAM.
Laptop OEMs didn't even want the Steam Deck chip or Broadwell-edram back then, much less a big expensive APU.
Much like how Apple sells 128 bit wide (mini and mba), 256 bit wide (m1/m2 pro), 512 bit wide (m1/m2 max), and 1024 bit wide (m1/m2 ultra).
I think the desktop/laptop vendors would have jumped at a nice iGPU/APU when they couldn't get normal GPUs.
Obviously AMD agrees, they shipped the PS5/XboxX and have a 256 bit wide APU planned for 2024. Just way later than I had hoped.
BTW, what is the usual width of memory buses in discrete GPUs?
In the previous gen even a relatively low end card like the 3060 Ti has a 256 bit wide memory interface. In the current generation the 4070 (which is a higher in the product stack) is 192 bits. The RTX 3080 (prev gen) is 320 bits wide the 4080 (current gen) is 256 bits.
Generally the trend is more cache, less width, and less bandwidth. Which is great ... for things that are cache friendly, but not everything is.
This is precisely why AMD, Intel, and Nvidia should think about making workstation-class machines with the lowest end of these - because until more people have one to play with, there won't be much to do with them.
Its a chicken and egg problem, I think. A big APU is so expensive that it doesn't really make sense without a very specific workload (like in a console), and the workloads dont really appear without the APUs.
Afaik, because the memory controller is part of the CPU, the CPU-RAM connection on the consumer chips is entirely passive - just copper traces on the motherboard with no ICs inbetween?
Perhaps that's not a fair comparison. From what I know AMD and NVIDIA use GPGPU cores (now with AI-focused instructions) plus separate AI-specific accelerator blocks. Conceptually, GPGPU + NPU on one die. NPUs can be much simpler than general-purpose GPUs. So AMD's driver and software stack likely needs to be an order of magnitude more complex than the NPU vendors' in order to accommodate other non-AI use cases. But to an end user it doesn't really matter why it sucks, only that it does.
Intel takes this approach too.
So the biggest HBM you can get is 24GB, and 8 of them is a reasonable max.
I believe the Apple M series is limited by package space? The LPDDR5(X?) bus is not physically/electrically limited to 192GB like the MI300.
It would be nice if this became on option for something like RunPod or Modal. Especially if it could be slightly cheaper than the Nvidia hardware somehow.
https://huggingface.co/TheBloke/falcon-40b-instruct-GPTQ
Large context size is becoming less of an issue now too.
But maybe it would be good for batched inference?
And for multiple services... Probably just better to run multiple cheap instances and/or load dynamically? The MI300 is super expensive.
In theory its more expensive to produce than an H100, but in practice... shrug.
But I was thinking 3090 or lesser tesla/Quadro instances would be sufficient.
Also ironically, this version of Falcon will require CUDA.
But it's just too much of investment for me on something that MAY work. I ended up just buying a 4080rtx
TBH the whole space is kinda a mess. Tons of optimizations (like most ML compilers), even on Nvidia cards, are left on the table because the SD UI devs just dont have the throughout or motivation to implement them.
At the other end, hardware makers, ml compiler devs, researchers and such are making quick demos, but are not making any integration attempts for popular frameworks.
There is no one in the middle, so we are stuck with PyTorch eager mode and a perception that it only works on big Nvidia GPUs.
But I don’t know whether it works with the more recent RDNA2 cards.
It "technically" works, but their own examples crash, you get a fraction of the performance you'd expect for the level of hardware you have, chicken&egg problem with little other software having good support of it, etc.
The problem with AMD is definitely not modern node access, but software, and with some investment it can probably change really fast.
Apple, on the other hand, is definitely not going to release standalone apple silicon products (GPUs or CPUs), so I don't really believe for Apple Silicon as an AI platform (apart from on-device inferrence).
I am skeptical too, but hardware design takes a looong time, and surely Apple sees the writing on the wall and the potential of their own hardware.
And they are optimizing more for power efficiency in multimedia workloads than raw AI throughout at any cost like the AI ASIC makers.