$95 AMD CPU Becomes 16GB GPU to Run AI Software
tomshardware.com
tomshardware.com
2s/iter? That is... Not great.
It is slower than CPU diffusion: https://github.com/rupeshs/fastsdcpu
Stable Diffusion in particular doesn't need much VRAM anyway. I get that many people are stuck on lower end computers, but ~4GB is not a totally unreasonable requirement.
VoltaML has some CPU optimized options for the torch.compile optimization.
Wow, did it really get that low? I played with SD during the initial days, and I know there was some strides then to get it running on 8gb cards and maybe squeaking through on 6gb with a smaller image size (like, 256x256px).
Can you generate reasonably large images on 4gb now, with just a higher wait time? I was never sure why it wasn't simple to just shuffle bits in and out of main RAM but I guess the model themselves are mostly big binary blobs that are hard to split up?
For SD 1.x and even 2.x, yes, in A1111 I ran both on a 4GB card before upgrading to 16GB 3080Ti laptop, and not even with --lowvram; reportedly, --lowvram will let you do at least 1.x on a 2GB card at normal (512x512) resolution directly.
> Can you generate reasonably large images on 4gb now, with just a higher wait time?
Yes, especially with tiled diffusion/multidiffusion.
Always has been. I can generate 768x768 with 3.9GB used, and I don't even use --lowvram or --medvram options in Automatic1111. But note that if you don't have an Intel CPU with onboard video, you will need to lower your resolution because because your OS is using a bunch of VRAM for the desktop already, more than the 0.1GB you left free.
SDXL is a total nonstarter as well on 8GB.
fine = requires --medvram or --lowvram in order to not run out of memory. The speeds above are with medvram, which is just about on the edge of what's possible. These options compromise speed but are still better than running on the CPU generally.
You are stuck at low res, but you can use a GAN to upscale the output.
Steam Deck is RDNA2 though, and is significantly faster. I don't think RDNA1 ever made it into an APU.
(74 days ago )
"AMD Ryzen APU turned into a 16GB VRAM GPU and it can run Stable Diffusion"
Would be interesting to see how this performed running Llama 2 7b.
Even then, interactive (batch=1) inferencing was still barely faster. I suspect memory bandwidth (I ran w/ 2 x DDR5-5600) is the main bottleneck. Just as a point of reference, the inferencing performance is roughly in the ballpark of (but slightly lower than) my 16GB M2 MBA.
Details: https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...
The Radeon 780M does 8.9 TFLOPS of FP32, 17.8 TFLOPS of FP16, so if 20 flops/byte applies, it has roughly 10X more compute than memory bandwidth for inferencing.
You're a bit apostrophe happy. Sure, as a matter of style and preference you can use apostrophes to pluralize short words / acronyms, but "it's" is a contraction for "it is".
An AMD Ryzen 2700X system for $1500? That might've been a reasonable price in 2018, but you couldn't reasonably spend that kind of money nowadays if you tried.
I'd love to see some more numbers showing implementations that specifically target certain platforms. For instance, you can run LLaMa on a Raspberry Pi, if you really wanted, with certain compromises. How would a fleet of small, inexpensive systems compare with a CPU-based and GPU-based systems?
It's interesting, all the same :)
It's true, but it seems orthographically inconsistent with other possessive forms; after all, "John's" can be a contraction for "John is," or it can be the possessive form of the name "John." By analogy, one might think that "it's" can also take on both meanings, but that's not the case.
Its even sadder that AMD/Intel have failed to take advantage of the price gouging. They could have released affordable 32GB-48GB cards with (afaik) no silicon changes.
The 7900 XTX is still somewhat reasonable where you can get it to work.
And the iGPUs are still pretty small.
We will see how Meteor Lake and Battlemage shake out, but I don't like the delay rumors.
If they intend to release an AI card, I bet they'll charge through the nose for it, just like Nvidia does.
But what they can do is double up the GDDR chips (which is what the 48GB pro cards do), kinda like running 2 DIMMS per channel on desktops. This requires a somewhat more expensive PCB and support from the memory controller, but AMD already supports this for sure, and I would be shocked if the Intel A770 die can't do it.
Impressive it works at all but the difference is night and day compared to top of the line (gamer card) Nvidia. Like SDXL <7 seconds a generation vs 60+ seconds.
(Assuming both machines are running the same model. I see a lotta crazy numbers thrown around based on comparing very different quantizations)
2 x 4090 is more like 38 tokens per second.
But it's harder than that. As the 2 x 4090 is using https://llm.mlc.ai/ which is highly optimized so perhaps the Mac Studio inference could be better when tuned.
Then a lot of production inference engines use continous batching, which means they perform really well when processing multiple prompts at a time.
As of the latest 8 bit cache commit, I can manage 2.6bpw with seemingly good quality.
For training, Apple has no CUDA, so you're pretty much dead in the water, unless you're specifically interested in e.g. local fine-tuning (for e.g. privacy or start-up strategy of "use the user's compute").
For inference, 4090 and H100 are not designed for throughput per $. Nvidia GPUs like the A40 or T40 are tuned for cloud inference, but at the end of the day your product is either running inference server-side (so play your own CAPEX vs OPEX, maybe even try Jetson) or user-side and so the user chooses the hardware not you.
I mean maybe there are Wallstreet analysts or people at Qualcomm/Broadcom who want an Apple vs Nvidia shootout to see who is getting the most of their silicon, but for the product space the comparison is basically irrelevant. Which is likely what Apple and Nvidia want right now.
It’s a bit weak for LLMs under near any circumstances but thinking it could probably handle two whisper cpp decodes over opencl and maybe some embeddings api stuff.
Is there something that prevents AI software from augmenting VRAM with system RAM? My impression was that for instance OpenGL doesn't even have a standard way to get VRAM usage because you're not supposed to care, stuff is just shuffled around in the background for you if VRAM runs out.
Depending on the type of model/usecase/how much you can fit in VRAM it very well may still be usable, however compared to fitting everything in vram it'll be abysmal for most any model type.
From my understanding the entire model must be read per iteration (or very close to all of it. This is why offloading any part of it impacts the performance so much.
Diffusion models depending on the sampler and some other stuff can potentially get away with only a few steps (usually 2-20 best case) and are relatively small (3B or less).
However LLM's are generally much bigger (they mostly start at 7B, there's a few exceptions depending on the use case), and they need an entire iteration per token which is only a few characters (think 2 to 3). So llm's especially are far more useful if you can keep them in vram.
As I understand, it adds complication, the points where it is practical without just getting into an horrible churn of shuffling are limited based on model architecture, and (as a consequence) implementations may not always been transferrable between models outside of a narrow family.
On the other hand, there are some frameworks with general support which eases this.
On the third hand, shuffling stuff back and forth between system RAM and VRAM has a performance cost in the best case.
So, yes, and no, and yes.
Most (that I've seen) consumer-focused frontends for local models (text generation, image generation, audio generation) do it, to a greater or lesser extent.
The base model implementations that are getting thrown out by researchers don't usually focus on that, and its generally implemented by third parties after the fact, sometimes with commonalities that can be applied to other models that are closely-enough related.
This includes hobby projects, tinkering out of curiosity, tasks performed for learning purposes (not necessarily educational) and more.
I guess this article just rubs my net negative lifestyle and feels like its content for contents sake.