It’d be short sighted to not have a wrapper around which AI to use, even without implementation for them. Like, having an abstract parent class with a Siri implementation of it would be baseline good software engineering.
People have speculated that with FPGAs or even an ASIC you could take some smaller models we have today, turn the chip into essentially that model-as-hardware, and get some really compelling performance/power-usage characteristics.
At that point running a GOOD quantized Q8 model on a phone is not beyond the realm.
That's what I was hoping for with Taalas, ultimately ending up with a microsd-ish card that's hot swappable intelligence.
That arrangement will change, somewhat, later today.