https://www.hardware-corner.net/guides/gpu-benchmark-large-l...
https://www.hardware-corner.net/guides/gpu-benchmark-large-l...
Power consumption is almost 10x smaller for apple.
Vram is more than 10x larger.
Price wise for running same size models apple is cheaper.
Upper limit (larger models, longer context) is far larger for apple (for nvidia you can easily put 2x cards, more than that it becomes whole complex setup no ordinary person can do).
Am I missing something or apple is simply currently better for local llms?
On apple silicons you can always use MoE models, which work beautifully. On RTX it's kind of waste to be honest to run MoE, you'd be better off running single, whole active model that fills available memory (with enough space for the context).
The OG DeepSeek models are hundreds of GB quantized, nobody is using RTX GPUs to run them anyway…
This just is just a single user chat experience benchmark.