(There is the Neural Engine which supports lower precision, but it limited in various ways.)
Regardless, the strides that Apple has been making are impressive.
(There is the Neural Engine which supports lower precision, but it limited in various ways.)
Regardless, the strides that Apple has been making are impressive.
Apple have an opportunity, if they 2-4x the memory on the entry level devices (not beyond the realms of possibility), to make local inference a thing available to all.
A lot of work is going on in 8-bit inference and even 4 bit inference. So, models that need 64 GB in FP32 can do with 16GB VRAM in FP8 or INT8, which is well within the realm of consumer NVIDIA cards. And the latest NVIDIA tensor cores will absolutely destroy Apple Silicon GPUs or the Neural Engine in 8 bit.
So, I don’t think it’s really a strong argument. And as someone who is a Mac user and a ML practitioner, I’d be very happy if they started supporting eGPUs again.
Apple Silicon has many strengths and the GPU core are fine for many ends, from games to graphics apps.
But let’s not pretend that Apple is beating NVIDIA at their own game (yet). That day might come, but currently it only leads to disappointed users in ML forums who were hyped into thinking that their vanilla M2 MacBook Airs can almost compete with a 4090 in training a deep transformer model. (Yes, that happens.)
Although quantisation has now lowered the required memory, I wouldn’t be surprised if it comes in handy again in the near future.
But today that’s quite a niche use case.
Then there's the power consumption difference to consider. This seems like one of those cases where benchmarks reveal only a fraction of the larger picture.