Which is neat, but it's not CUDA. It's an application-specific accelerator good at a small subset of operations, controlled by a high-level library the industry is unfamiliar with and too underpowered to run LLMs or image generators. The NPU is a novelty, and today's presentation more-or-less confirmed how useless it is for rich local-only operations.
> Apple also has the means to build data center training hardware using apple silicon if they want to do so.
They could, but that's not a competitor against an NVL72 with hundreds of terabytes of unified GPU memory. And then they would need a CUDA competitor, which could either mean reviving OpenCL's rotting corpse, adopting Tensorflow/Pytorch like a sane and well-reasoned company, or reinventing the wheel with an extra library/Accelerate Framework/MPS solution that nobody knows about and has to convert models to use.
So they can make servers, but Xserve showed us pretty clearly that you can lead a sysadmin to MacOS but you can't make them use it.
> they could also start to supply them with cloud inference hardware and strongarm them into only using apple servers to serve iOS requests.
I wonder how much money they would lose doing that, over just using the industry-standard Nvidia servers. Once you factor in the margins they would have made selling those chips as consumer systems, it's probably in the tens-of-millions.