Peak theoretical throughput for the GPUs you find in ARM SoCs is quite good compared to the power draw, but you will not get peak throughput for workloads designed for Nvidia and AMD GPUs.
Peak theoretical throughput for the GPUs you find in ARM SoCs is quite good compared to the power draw, but you will not get peak throughput for workloads designed for Nvidia and AMD GPUs.
Ray tracing outside of Nvidia is a disaster all round, so yeah, nobody is competing on that front.
But among people running LLMs outside of the data centre, Apple's unified memory together with a good-enough GPU has attracted quite a bit of attention. If you've got the cash, you can get a Mac Studio with 512GB of unified memory. So there's one workload where apple silicon gives nvidia a run for their money.
[1]: https://www.notebookcheck.net/Cyberpunk-AC-Shadows-on-Apple-...
It's not just power budget, the desktop part has more of everything, and in this case the 4070 mobile vs desktop turns out to be a 30-40% difference[1] in games.
Now I don't have a mac so if you meant "2-5x" when you said "much faster" well thdn yea, that 40% difference isn't enough to overcome that.
[1]: https://nanoreview.net/en/gpu-compare/geforce-rtx-4070-mobil...
That should be feasible, no?