I don't think it's a very huge or exploitable niche for two reasons:
- Android, Windows, Linux and MacOS can already run local and private models just fine. Getting something product-ready for iPhone is a game of catch-up, and probably a losing battle if Apple insists on making you use their AI assistant over competing options.
- The software side needs more development. The current SOTA inferencing techniques for Apple Silicon in llama.cpp are cobbled together with Metal shaders and NEON instructions, neither of which are ideal or specific to Apple hardware. If Apple wants to differentiate their silicon from the 2016 Macbooks running LLaMA with AVX, then they have to develop CoreML's API further.