AMD Ryzen APU turned into a 16GB VRAM GPU and it can run Stable Diffusion
old.reddit.com
old.reddit.com
That might be a fine tradeoff for some kernels. But my understanding is that Stable Diffusion is very VRAM-bandwidth heavy and actually benefits from the higher-speed GDDR6 or HBM RAM on a proper high-end GPU.
edit: comment from the reddit thread:
> LLM isn't all that great as it is primarily memory bandwidth bound, ie almost no difference from a CPU if your memory bw is mere 12/25Gb/s. SD needs far more compute for inference - APU with slow memory helps.
- https://old.reddit.com/r/Amd/comments/15t0lsm/i_turned_a_95_...
2x sticks of 32GB/s RAM, properly configured, will run at 64GB/s of bandwidth. Modern servers are quad, hex, or oct-channel (4x, 6x, or 8x parallel sticks of RAM) in practice. Or even more (ex: Intel Xeon Platinums are 6x channel per socket, so an 8x socket 8x CPU Xeon Platinum will be like 48x DDR5 parallel RAM sticks).
----------
PCIe x16 by the way, is 16x parallel lanes of PCIe. All the parallelism is already innate in modern systems.
---------
L3 cache is TB/s bandwidth IIRC. CPUs inside of CPU-space will automatically be caching a lot of those RAM commands, so you can go above the RAM-bandwidth limitations in practice, though it depends on how your code accesses RAM.
GPUs have very small caches, and have higher latency to those caches. Instead, GPUs rely upon register-space and extreme amounts of SMT/wavefronts to kinda-sorta hyperthread their cores to hide all that latency.
https://www.ebay.com.au/itm/126046615075
That's $250 in Australian dollars though, which is about US$160. I'm not affiliated with that seller btw, I just remembered the search result from looking a while back. :)
Here’s some photos https://aus.social/@s_mcleod/110841559904867676
Is that $250 Australian, or $250 US?
And of course you have to be happy with the additional 30 W idle power consumption when you’re not using it if it’s in your home server, which as I’m sure you’ll appreciate in Australia can be quite expensive now!
If they cost more (say $350+) I probably go get a slightly more expensive, but much more modern RTX.
If AMD improves their speed with this configuration on later APUs though - that could really hurt Nvidia.
And most people playing with AI don't have an expensive GPU with 12+ GB VRAM.
That’s using a 4600G.
These chips don’t have a socket and were designed for laptops. However, they have up to 54W TDP which is not quite a laptop’s territory. Luckily, there’re mini-PCs on the market with them. The form factor is similar to Intel NUC or Mac Mini. An example is Minisforum UM790 Pro (disclaimer: I have never used one, so far only read a review).
The integrated Radeon 780M GPU includes 12 compute units of RDNA3, peak FP32 performance is about 9 TFlops, peak FP16 about 18 TFlops. The CPU supports two channels of DDR5-5600, a properly built computer has 138 GB/second memory bandwidth.
I have tried to find an authoritative source but failed, because DDR5 spec on jedec.org is sold for $370 :-(
I have a 7940HS on my desk now and the Memtest reports the same number.
Apple Silicon has extremely high memory bandwidth which is why it performs so well with LLMs.
DDR4-4800 exists. 76.8GB/s. You can also get a Ryzen 7000 series for around $200 that can use DDR5-8000, which is 128GB/s. By contrast, the M1 is 68GB/s and the M2 is 100GB/s. (The Pro and Max are more, but they're also solidly in "buy a GPU" price range.)
Also, when doing a test like this it's important to compare same bit depths so fp32 on both.
But if something fits inside of CPU-cache space, its probably faster on CPU. (Intel is upto 2MB L2 cache on some systems, AMD is up to 192MB L3 cache on some systems).
But if something is outside of CPU-cache space but fits inside of GPU-VRAM space, its probably faster on GPU. (Better to be bandwidth-limited at 500GB/s GPU-VRAM speeds than 50GB/s DDR5 speed). Ex: 16GBs GPU-VRAM is back into GPU-sized solutions to a problem.
Then if something fits inside of CPU-RAM space, which is like 2TB in practice (!!!!), then CPUs are probably faster. Because there's no point taking a 2TB DDR5 RAM limited problem, passing some of that data to 16GB or 96GB of GPU-VRAM, and then waiting for DDR5 RAM anyway.
------------
Not that I've tested out this theory. But it'd make sense in my brain at least, lol.
Another thing, it’s hard to write CPU code which saturates memory throughput instead of stalling on memory latency. In GPUs, the issue is bypassed on architecture level. They enjoy a high degree of concurrency; GPU cores simply switch to another thread instead of waiting. The CPU cores only run 2 threads each due to hyper-threading, not enough threads to efficiently hide memory latency the way GPUs are doing.
ROCm and PyTorch recognize the GPU, as in, `torch.cuda.is_available` returns true, but I haven't actually run any models yet.
The maximum I can select on my 32GB RAM laptop is 8GB. Sounds like there are desktop AM4 mainboards where you can go up to 64GB.
One interesting thing I have seen being reported by `glxinfo` is auxiliary memory, which in my case is another 12GB, for a total of 20GB reported total available memory GPU memory. Unclear if this could be used by ROCm/PyTorch.
The biggest issue that keeps me from recommending it to anyone is that it is a windows 11 machine. That is both its strongest and weakest point. Turn it on for the first time, windows setup on a tiny screen with no physical mouse keyboard. And then there is messing with game graphics settings to eek out the performance you want. SD slot is complete garbage. Most likely a hardware issue that will not be able to be fixed. It stops reading cards and in some cases, destroys the card. It is fairly easy to upgrade the internal SSD though. ASUS has also been pretty quiet on graphics driver updates.
As long as you are cool with that stuff, its an awesome machine. Steamdeck has more of a 'just works' kinda console feel and kinda outdated hardware wise. Lenovo Go images just leaked today.
~2016 I bought a Lenovo ThinkPad E465 sporting an AMD Carrizo APU specifically to take advantage of HSA. It seemed like the feature, or at least the toolchain to take advantage of it, never really materialized. I'm glad at least someone else remembers it.
https://gpuopen.com/ags-sdk-5-4-improves-handling-video-memo...
Which says: "For APUs, this distinction is important as all memory is shared memory, with an OS typically budgeting half of the remaining total memory for graphics after the operating system fulfils its functional needs. As a result, the traditional queries to Dedicated Video Memory in these platforms will only return the dedicated carveout – and often represent a fraction of what is actually available for graphics. Most of the available graphics budget will actually come in the form of shared memory which is carefully OS-managed for performance."
The implication seems to be that you can have an arbitrary amount of graphics RAM, which would be appealing for AI use cases, even though the GPU itself is relatively underpowered.
Still, the question remains open, how to precisely control APU/GPU memory allocation on Linux and what is the limitations?
There is videos on YouTube of it with 2TB of flash.
-----begin copy paste-----
The 4600G is currently selling at price of $95. It includes a 6-core CPU and 7-core GPU. 5600G is also inexpensive - around $130 with better CPU but the same GPU as 4600G.
It can be turned into a 16GB VRAM GPU under Linux and works similar to AMD discrete GPU such as 5700XT, 6700XT, .... It thus supports AMD software stack: ROCm. Thus it supports Pytorch, Tensorflow. You can run most of the AI applications.
16GB VRAM is also a big deal, as it beats most of discrete GPU. Even those GPU has better computing power, they will get out of memory errors if application requires 12 or more GB of VRAM. Although the speed is an issue, it's better than out of memory errors.
For stable diffusion, it can generate a 50 steps 512x512 image around 1 minute and 50 seconds. This is better than some high end CPUs.
5600G was a very popular product, so if you have one, I encourage you to test it. I made some videos tutorials for it. Please search tech-practice9805 for on Youtube and subscribe to the channel for future contents. Or see the video links in Comments.
Please also follow me on X: https://twitter.com/TechPractice1 Thanks for reading!
https://old.reddit.com/r/Amd/comments/15t0lsm/i_turned_a_95_...
ROCm has kinda worked for some APUs in 5.x for a while now. As much as things are expected to work on AMD hardware anyway.
So basically install ROCm 5.5 and check if rocminfo lists your APU as device. Assign more VRAM if possible in your Bios.
There are really no fundamental secret tricks involved. I run pytorch on a 5600G. It's not great and breaks all the time, but that's not really APU specific.