Any PC regardless of CPU architecture with 16GB+ can do this, right?
Not really, no. You want to use the GPU not the CPU. Macs are neat here since they can use the shared memory with a rather high bandwidth. So even if the GPU is much slower, the ram is much worse than proper vram, and ridiculously overpriced for that,... often the bottlenecks are ram amount and bandwidth.
Having run ollama on CPU: Yes, it's just slower. Not even intolerably slow IMO, though I used small models and don't mind some turnaround time.
I've seen llamafile go about 10x faster on CPU if you try it.
Really! Same model and everything? I guess I need to go benchmark them - 10x would legit obviate GPU for me
You're supposed to act impressed that they recommended an Apple product.
And yes, but there's no lower limit on the memory; it's entirely dependent on the model or kernel size.
I only have some light use cases, so I use a cheap laptop (<$250) with a ryzen APU 8gb soldered/shared ram. Then added a 16gb ram stick, booted off a usb bios from github, increased uma buffer to 8gb. I had stable diffusion working, it was slow, but I'm pretty sure its faster than cpu/ram (2 min for 512x512 20-25 step).
I've run it on a Pi 5 with 8Gb, and get about a token a second
M-series are a LOT faster :)
Even my M1/16GB gets decent speeds. 7+ tokens/second with llama3