I'm a big fan of two-tier systems where the 90% case is taken on by something cheap, and renting the premium tier. Works for cars, server hosting, dining ...
I'm a big fan of two-tier systems where the 90% case is taken on by something cheap, and renting the premium tier. Works for cars, server hosting, dining ...
Even assuming you generate 100 tok/second, 24 hr, that's only 8 mil tokens. About $5 for Gemini 2.5 Flash or $50 for 2.5 Pro. But you will not be able to run it 24/7 locally - sleep, thinking between prompts, etc
Even at continuous usage (24 hr/day), it's still cheaper to just pay Google for Gemini 2.5 Flash tokens for half a year.
And this is NOT taking into account that mistral/devstral/qwen3 are vastly inferior models.
How can Google be so cheap? TPUs, batching multiple requests, economy of scale, loss-leader, ...
Running stuff locally makes sense for fun, as backup for loss of connectivity, if you want to try uncensored models, but it doesn't make sense economically.
For investigating concepts, can AI for this at all? Subscriptions will be a better approach.
Local models are underutilized in comparison to the promise they have. Imagine browser-agnostic universal AdBlock and that's just the surface.
If I worked on a codebase I couldn’t send to the hosted tools I’d go this route in a heartbeat.
I think if you’re comparing it to Claude max it’s ok, but the payback period on GitHub copilot for example is almost 7 years.
This is probably the easiest problem to solve, isn't it? Most tools offer a n API endpoint setup that you can use locally and also expose on a local network easily.