2 years ago you had downloaded onto your laptop an effective and useful summary of all of the information on the Internet, that could be used to generate computer programs in an arbitrarily selected programming language.
I wrote a post about it: Your toaster will know mesopotamian history because it’s more expensive not too.
https://wanderingstan.com/2026-03-01/your-toaster-will-know-...
I'm sure people would get a cheaper toaster in exchange of an ad being burned in your bread.
And as other commenter pointed out, a smart toaster with ads or data collection can be subsidized and thus be more profitable. (Oh what a world we're headed for!)
In any case, I think the LLM-everywhere thesis holds even strong for even moderate-complexity devices like power plugs, microwaves, and mobile phones.
But will it know the difference between too and to?
Ask your local model a verifiable question - for example a list of tallest buildings in Europe. I did it with Gemma on my laptop, and after the top 3 they were all fake. I just tried that again with Gemma-4 on my iphone, and it did even worse - the 3 tallest buildings in Europe are apparently the Burj Khalifa, the Torre Glories and the Shanghai Tower.
I wouldn't call that effective compression of information.
But what you can do with local models is give them actual data and tools to search it. Download a copy of Wikipedia locally, give the agent a way to search it and BOOM accurate information without an internet connection.
Also "small enough to live on disk" is a bit vague, especially when models get super stupid super fast when you get to the smaller size. At that point they're just basically 40k servitors that can use tools and nothing much.