When I was a SWE, I've done migrations between bare metal to AWS to GCP and then both, and then all plus Azure..
It is the cost of doing business. You pick what works best at a price that is optimized for your business needs now. You have a war chest so when they start becoming assholes, you have leverage to fight back or pivot.
I'd rather spend $200-400/mo to unblock myself NOW than do something dumb with 5 or even 100 tokens per second output that isn't actually that good as what the current providers offer. I'm going through millions of tokens a day.. I couldn't do that with local "RIGHT NOW" (<--- important)
Unless the actual race is to create an AI Employee that operates and can deliver work without constant supervision. At that level, of course it would be cheaper to pay $2,000 a month straight to Open AI vs hiring a junior SWE
Or is this the case of every HN discussion where what you do is “actual work” and what other people do is “toying around?”
I agree with this... for now. But the hosted commercial models aren't widening the gap as far as I can tell, if anything it appears to be narrowing.
And if the relative delta doesn't increase somehow I don't see any way in which the "AI race" doesn't end in a situation where locally run LLMs on relatively cheap hardware end up being good enough for virtually everyone.
Which in many ways is the best possible outcome, except for the likely severe economic effects when the bubble bursts on the commercial AI side.
I don't know what will happen tomorrow, let alone 1 year or more down the line.
If it ever becomes economical to run and maintain bare metal GPU compute to run LLMs, then that's what will need to be done..