Boosting Computational Fluid Dynamics Performance with AMD MI300X
rocm.blogs.amd.com
rocm.blogs.amd.com
*I realize enterprise “GPUs” are difficult to purchase as an individual whether they’re AMD or Nvidia, but AMD might be well-served to mimic their RX 480 strategy (“make a competitive mid-range GPU, distribute it through several board partners, and sell it at rock-bottom prices to get it to as many customers as possible”). If there’s a compelling reason to use AMD hardware over Nvidia, the software support will come. As an extreme example, if one could purchase an MI300X for $50 a pop, surely FAANG and others would invest time and effort into the software side to leverage the cost savings vs Nvidia, who is unquestionably price-gouging due to their monopolistic hold on the enterprise GPU market.
For example this blog post, about how great MI300X is. Really, what do I care -- I'm not a billionaire.
The whole AMD AI/ML strategy feels like this - prioritize short term profits and completely shoot themselves in the foot in the long term.
That's what the stock market rewards.
DirectX shaders however were already ready for Wave32, and other architectural changes that RDNA had. In fact, RDNA was basically AMD changing their architecture to be more "NVidia-like" on many regards (32-wide execute being the most noticeable).
CDNA existed because HPC has $Billion+ contracts with code written for Wave64 and still needing ROCm support. That means staying on the older GCN-like architecture and continuing to support say, DPP instructions or other obscure features of GCN.
---------
Remember how long it took for RDNA to get ROCm support? Did you want to screw the HPC customers for that whole time?
Splitting the two architectures, focusing ROCm on HPC (where the money was in 2018 for GPU Compute research dollars), and focusing on better video game performance for RDNA (where money is for video game / consumer cards) just makes sense.
Not really.
Wave64 on CDNA is provably more throughput. But with most video game code written for NVidia's Wave32, RDNA being reworked to be more NVidia-like and Wave32 is how you reach better practical video game performance.
HPC will prefer the wider execute, 64-bit execution, and other benefits.
Video Gamers will prefer massive amounts of 32MB+ of "Infinity cache", which is used in practice for all kinds of screen-space calculations. But this would NEVER be used for fluid dynamics.
CDNA executes 64-threads per compute unit per clock tick. RDNA only executes 32-threads. CDNA is smaller, more efficient, more parallel and much higher compute than RDNA.
Furthermore, all ROCm code from GCN (and older) was on Wave64, because historically AMD's architecture from 2010 through 2020 was Wave64. RDNA changed to Wave32 so that they can match NVidia and have slightly better latency characteristics (at the cost of bandwidth).
CDNA has more compute bandwidth and parallelism. RDNA is narrower, faster latency and less parallelism. Building a GPU out of 2048-bit compute (aka: 64-lanes x 32-bit wide/CDNA) is always going to be more bandwidth than 1024-bit compute (aka: 32-lanes x 32-bit wide) like RDNA.
If you actually were using both, you'd know that CDNA was the only supported platform on ROCm for what felt like an eternity. That's because CDNA was designed to be as similar to GCN so that ROCm could support it easier.
--------
What I'm saying is that today, now that ROCm works on RDNA and CDNA, the two architectures can finally be unified into UDNA. And everyone should be happy with the state of code moving forward.
AMD needs to invest a Fsck load of money in software... Until then they can have the greatest compute cards in the world.but it will.mean nothing
My favorite one of many examples of this is when Nvidia said “okay, fine, we’ll allow customers to enable G-Sync (adaptive refresh) on any display, but only 1000 series and newer GPUs!” No hardware limitation for this, they just didn’t want to give any 900-series holdouts a reason not to upgrade.
Then there’s the arbitrary lock on using consumer Nvidia GPUs in passthrough on a level 1 hypervisor. Sometimes there are unofficial workarounds but why do I have to go through that hassle at all? Ain’t nobody realistically going to buy consumer grade GPUs to throw in datacenter racks… but even if they were, so what? It’s their hardware, they bought it, let them use it. No reason whatsoever to block passthrough on consumer cards except $$$.
This is what I saw as well. As a developer, I wanted access to enterprise HPC compute, but I'm also not going to do a PhD just to play around with these things. So, I got funding, started a business and bought 8 of them as a PoC test. We got customers, we got more funding, got a real datacenter, we bought 128 more. Crawl, walk, run.
You can now rent them by the minute from us for a few bucks an hour. Currently limited to docker containers for individual GPUs, but you can get a full bare metal 8x box too (with BIOS too!). Support for VM's is coming. If you want multiple boxes, we have the full 8x400G NICs too. The boxes are fully loaded with tons of enterprise NVMe, RAM and top core/clock Intel CPUs (not AMD cause Dell didn't have that as a solution).
Our model is to follow AMD's roadmap and buy/release their products as they come. We're currently debating the 325x and looking forward to / planning for the 355x.
Despite your desire, it will be a long time before there is a consumer version of these things. Especially as they move to more and more complex deployments. Look at the NV72 and the requirements around that... we can all guess where AMD is going. DC rails in the racks, DLC cooling, massive power requirements. It is only getting more and more capex/opex intensive.
Let's also not forget that AMD is really just a hardware manufacturer. When you buy a RX480 (I had 130,000 of these previously), it was from an OEM, like Sapphire, that could handle all the end user support.
This is why the whole NeoCloud industry has sprung up. Large clouds can only handle this pace by selling thousands at a time in multi-year contracts. We are taking the long tail and built a business around that. Short of doing everything we are doing yourself (which trust me, is not easy), your best bet is to work with companies like mine to get you access to this gear.
You can now rent them by the minute from us for a few bucks an hour. Currently limited to docker containers for individual GPUs, but you can get a full bare metal 8x box too. Support for VM's is coming. If you want multiple boxes, we have the full 8x400G NICs too. The boxes are fully loaded with tons of enterprise NVMe, RAM and top core/clock Intel CPUs (not AMD cause Dell didn't have that as a solution).
Meanwhile I can’t even find an MI300X on eBay. I can at least poach enterprise Nvidia GPUs like the A100 on eBay. This tells me AMD’s shipping far fewer units and therefore enterprise GPUs aren’t doing much for their balance sheet (though I’d have to look at their quarterly and annual reports to know for certain). To me this strengthens the case for selling to individuals/startups, and at prices that offset the risk of picking AMD over Nvidia and potentially running into software shortcomings.
I’m set with two RTX 3090s at the moment, but it’s very neat that you’ve been able to bootstrap essentially a cloud service provider in the age of AWS, Azure and GCP (and DO and Vultr and Linode et al).
There absolutely is. The current form factor is not standard PCIe. It is a OAM/UBB board that is custom designed by AMD to support Infinity Link. It only comes in an 8x configuration. Now, you're asking for a totally different design and that requires a huge investment that would take away from their existing focus on enterprise.
> Meanwhile I can’t even find an MI300X on eBay.
They should have considered the RX 480 approach for the Instinct accelerators.
So, on one hand, people want them to compete with Nvidia, but on the other hand we want them to ignore the market that is actually going to make them money. As much as it would be nice to have, we can't have it both ways.
The middle ground is to rent a single one from us billed by the minute (or another neocloud). We handle all the detailed problems for you (don't forget the massive upfront capex spend), and you get to build your products/companies on that. Once you grow to the point of being able to buy your own equipment, we can even help you deploy it.
Some of this is lack of groundwork/engineering by packages or system administrators, but it seems a decent amount is the relative lack of effort by AMD to make things work well OOTB.
It won't, not in any way that will make AMD approximately competitive with Nvidia.
AMD, unlike Nvidia, seems unable to prioritize developers. Here's a summary of last week's charlie-fox when the TinyGrad team attempted to get 2 MI300s for on-premises testing and was rebuffed by an AMD representative. https://x.com/dehypokriet/status/1879974587082912235
It will get better, but Nvidia seems to be creating CUDA libraries for all kinds of applications, so the moat is constantly widening/deepening.
Anush (AMD VP of AI software) has had a fire lit under his butt after the recent SemiAnalysis article [1] and is actively taking feedback on improving the experience. If you have specific things you'd like to see, I'm more than happy to forward them onto him (contact in my profile).
[0] https://rocm.docs.amd.com/en/latest/
[1] https://semianalysis.com/2024/12/22/mi300x-vs-h100-vs-h200-b...
I would not be surprised at all if these benchmarks ran faster if you removed the GPUs completely.
I've never seen a non-reactive incompressible flow simulation get substantial speedup on GPUs. There are well understood fundamental reasons why this is the case.
Now granted, the flops to byte ratio for this program might be better than an avg fluid simulator. Also, our performance tanked when we moved to multi-node system. But I am aware of underlying reasons behind the scalibility issues and they don't feel like problems that can't be overcome.
Especially if you do the comparison on equivalent cost basis, i.e. "what is the walltime difference if I run on a $60k all-CPU cluster versus a $60k GPU cluster". Or in terms of cloud compute cost / HPC allocation spend.