I can reasonably run (quantized) Mistral-7B on a 16GB machine without GPU, using ollama. Are you sure it isn't a configuration error or bug?
Even with 12 threads of my 5900X (I've tried using the full 24 SMT - that doesn't really seem to help) with the dolphin-2.5-mixtral-8x7b.Q5_K_M model, my MacBook Pro is around 5-6x faster in terms of tokens per second...
But whatever it is, it's great, and I hope that Intel and AMD will catch up.
AMD has had the APUs for awhile but I think they aren't at the same level at all as the new Mac acceleration.