Using cloud infrastructure should help with this issue. It may cost much more per run but money can be saved if usage is intermittent.
How are HN users handling this?
Using cloud infrastructure should help with this issue. It may cost much more per run but money can be saved if usage is intermittent.
How are HN users handling this?
Plenty of people in tech earn enough to support a family and drive a fancy car, but choose not to. A used RTX 3090 isn't cheap, but you can afford a lot of $1000 GPUs if you don't buy that $40k car.
Other options include only running the smaller LLMs; buying dated cards and praying you can get the drivers to work; or just using hosted LLMs like normal people.
To your point about cloud models, these are really quite cheap these days, especially for inference. If you're just doing conversation or tool use, you're unlikely to spend more than the cost of a local server, and the price per token is a race to the bottom.
If you're doing training or processing a ton of documents for RAG setups, you can run these in batches locally overnight and let them take as long as they need, only paying for power. Then you can use cloud services on the resulting model or RAG for quick and cheap inference.
I’m not saying that dumping $10k into rapidly depreciating local hardware is the more economical choice, just that people often discount the likelihood and cost of making mistakes in the cloud during their evaluations and the time investment required to ensure you have the correct safeguards in-place.
> How are HN users handling this? I’m working on a startup for end-to-end confidential AI using secure enclaves in the cloud (think of it like extending a local+private setup to the cloud with verifiable security guarantees). Live demo with DeepSeek 70B: chat.tinfoil.sh
I'm about to plunge in as others have to get my own homelab running the current crop of models. I think there's no time like the present.
There's a healthy secondary market for GPUs.
I'm seeing 24GB M40 cards for $200, 24GB K80 cards for $40 on eBay.
* Cheap
* Fast
* Decent amount of RAM
Pick two.
These old GPUs are as cheap as they are because they don’t perform well.
So Cheap and Decent amount of RAM work for me.
Combine the best of both worlds. I have a local assistant (communicate via Telegram) that handles tool-calling and basic calendar/todo management (running on a RTX 3090ti), but for more complicated stuff, it can call out to more advanced models (currently using OpenAI APIs for this) granted the request itself doesn't involve personal data, then it flat out refuses, for better or worse.