Also local models are close in capabilities now but who knows in a few years what that'll look like.
Yes, and the flow of future models may dry up, but the current local models we'll have forever.
For me the switching point will probably be when they (AI companies) start the big rug pull. By then my hope is self hosting will be cheaper, better, easier, etc.
I do use the Gemini assistant that came with this Android, in the same cases and with the same caveats with which I use Siri's fallback web search. As a synthesist of web search results, an LLM isn't half bad, when it doesn't come as a surprise to be hearing from one at least.
You can buy a used recent PC for a hundred or two, cram it full of memory, and then run a very advanced model. Slowly. But if you are planning to run an agent while you sleep and then review the work in the morning, do you really care if the run time is 4 hours instead of 40 seconds? Most of the time, no. Sometimes, yes.
Some people overload "local" a bit to mean you are hosting the model - whether it's on your computer, on your rack, or on your hetzner instance etc.
But I think parent is referring to the open/static aspect of the models.
If it's hosted by a generic model provider that is serving many users in parallel to reduce the end cost to the user, it's also theoretically a static version of the model... But I could see ad-supported fine tunes being a real problem.
Funny enough the mac has almost the same processor as my iPhone 16 Pro, so its just a RAM constraint, and of course PrivateLLM does not let you host an API.
An M4 Pro would do much better do to the increase in RAM and GPU size.
This has been the hardest thing for me to learn and since everything's evolving so quick, what's recommended one week might not be the next.
Funny enough the mac has almost the same processor as my iPhone 16 Pro, so its just a RAM constraint, and of course PrivateLLM does not let you host an API.