1. Apple doesn't provide public interfaces to the ANE or the AMX modules, making it difficult to fully leverage the hardware, even from PyTorch's MPS runtime or Apple's "Tensorflow for macOS" package
2. Even if we could fully leverage the onboard hardware, even the M1 Ultra isn't going to outpace top-of-the-line consumer-grade Nvidia GPUs (3090Ti/4090) because it doesn't have the compute.
The upcoming Mac Pro is rumored to be preserving the configurability of a workstation [1]. If that's the case, then there's a (slim) chance we might see Apple Silicon GPUs in the future, which would then potentially make it possible to build an Apple Silicon Machine that could compete with an Nvidia workstation.
At the end of the day, Nvidia is winning the software war with CUDA. There are far, far more resources which enable software developers to write compute intensive code which runs on Nvidia hardware than any other GPU compute ecosystem out there.
Apple's Metal API, Intel's SYCL, and AMD's ROCm/HIP platform are closing the gap each year, but their success in the ML space is dependent upon how many people they can peel away from the Nvidia Hegemony.
1:https://www.bloomberg.com/news/newsletters/2022-12-18/when-w...
Despite Apple engineers contributing hardware support to Blender: https://9to5mac.com/2022/03/09/blender-3-1-update-adds-metal...
So looks like NVidia (and AMD to some extent) are winning the 3D usage war as well for now.
This isn't really an Apple specific problem - GPUs with low power budgets are never going to beat GPUs intended to permanently run off of wall power.
(I know there's stuff like the Mac Studio and iMac, but those are effectively still just laptop guts in a (compact) desktop form factor rather than a system designed from the ground up for high performance workstation type use)
I'd love to see a dedicated PCIe GPU developed by Apple with their knowledge from designing the m1/m2, but it's not really their style.
But... the GTX by itself has a 145W TDP, whereas the whole laptop has a TDP of 20W!
This is super impressive imho, but of course not in the raw power department.
Would be very interesting to retest with the M2 and also in the months/years to come as the software reaches the level of optimisations we see on the PC side.
[1] https://wandb.ai/tcapelle/apple_m1_pro/reports/Deep-Learning...
5950X CPU ($500) with 128GB of memory ($400).
AFAIK it flat out won't work with the DL frameworks.
If I'm mistaken please do speak up.
Edit: Thank you JonathanFly, I didn't know this!
This can work in some cases (years ago I certainly did it for CNNs - was slow but you are fine tuning so anything is an improvement) but I don't know how viable it would be on a transformer.