For some tasks, yes. For most of my deeper work they're not even close to my subscriptions.
> It takes less time for local model to take the first action on your task than it does for Claude to validate your login, put you into queue and start issuing the commands.
I have some decent LLM hardware here and I strongly disagree with this. Claude responds quickly. Using Fable or Opus it will deliver a working result faster than my local models because it gets there in fewer tokens. That's just how it is.
> During the winter time the GPU also doubles as a 300W in-house heater.
This is a curse in the summer. I'm feeling it right now.
Just a reminder that heat pumps can consume 300W of electricity to provide 1200W of heat.
(Tangent: I wonder why we haven't seen deployment of organic Rankin cycle generators in AI data centers, the exhaust temperature should be compatible and that could yield a 10-20% energy bill saving).
https://eu-mayors.ec.europa.eu/en/news/stockholm-sweden-heat...
I'm talking about making electricity back from the heat (using a low-temp thermodynamic cycle). It has a low yield (due to the low input temperature) but it's usually economically viable when using heat that would end up in the heavens anyway.
Let's say your incoming water temperature is 18C and you want it preheated to 50C, which is 32C degree differential, which means you'll need 32 * 1.5 * 160 = 7680 Wh, or 25 hours straight to heat the buffer tank from scratch.
You'll need to purchase a small water-to-water heat exchanger ($50), two pumps ($100 each), a power supply for said pumps, hose and/or copper pipe and fittings, and various other sundries, plus the cost of a buffer tank ($600ish), so figure all in roughly $1000, plus the cost of electricity to run the pumps.
At $0.22/kWh you're saving roughly $450 a year with this setup in foregone water heating, but because it's not 100% efficient you're spending $500 in electricity to run your GPU 24/7/365, and that's the maximum you can possibly save with the above assumptions. Scale up for more GPUs and down accordingly for less usage as you see fit.
Alternatively, use air as the heat conductor by placing the GPU laden machine in the same space as a hybrid heat pump water heater.
Because there was never a long term plan for AI data centers. It's an AI market capture and cash grab scheme that ends when local models eat their lunch.
Did the rise personal computing make data centers and super computers obsolete?
Any advance in inference that allows local models to do the job will also benefit hyperscalers. Imagine the sheer amount of compute they could throw at problems if each 32GB of VRAM was enough for frontier reasoning.
But yes
There could be a service that works in reverse where if someone needs a heater for a few months, they could rent a portable server (e.g. using older repurposed GPUs) with a built-in 5G modem that would run inference on LLM queries. As an incentive perhaps renting itself could be free (or you could earn money?), but you'd still have to pay your electricity bill.
I should only have to pay 1/3 of what it adds to my bill, given that it's 1/3 as efficient as a heat pump. And that coefficient is to be adjusted as outside temp changes (assuming air source heat pump; water source would have a stable COP).
This is only true if your local model is already resident in RAM / VRAM.
Please correct me if I am wrong.
A bit slow for agentic coding of course but fine for any chatbot use-case.