When you say it can run on consumer gpus, do you mean pretty much just the 4090/3090 or can it run on lesser cards?
I was surprised by how fast it runs on an M2 MBP + llama.cpp; Way way faster than ChatGPT, and that's not even using the Apple neural engine.
It's more than fast enough for my experiments and the laptop doesn't seem to break a sweat.