I hope this is the future. Offline, small ML models, running inference on ubiquitous, inexpensive hardware. Models that are easy to integrate into other things, into devices and apps, and even to drive from other models maybe.
Such hardware is not general-purpose, and upgrading the model would not be possible, but there's plenty of use-cases where this is reasonable.
The tech is still public and the research is available
This way they could offload as much of the "LLM" work on a device that lives in the home, all family linked phones and devices could use it for local inference.
It's way overpowered as is anyway, why not use it for something useful.