There are a lot of problems here though. First of all being that inferencing isn't hard to do - iPhones were capable of running LLMs before LLaMa and even before it was accelerated.
Anyone can inference a model if they have enough memory, I think Nvidia is banking on that part.
Then there's the issue of model size. You can fit some pruned models on an iPhone, but it's safe to say the majority of research and development is going to happen on easily provisionable hardware running something standard like Linux or FreeBSD.
And all this is ignoring the little things, too; training will still happen in-server, and the CDN required to distribute these models to a hundred million iPhone users is not priced attractively. I stand by what I said - Apple forced themselves into a different lane, and now Nvidia is taking advantage of it. Unless they intend to reverse their stance on FOSS and patch up their burned bridges with the community, Apple will get booted out of the datacenter like they did with Xserve.
I'm not against a decent Nvidia competitor (AMD is amazing) but the game is on lock right now. It would take a fundamental shift in computing to unseat them, and AI is the shift Nvidia's prepared for.