I respect Anthropic (and OpenAI to a lesser extent) but I’m not going to play these games.
I respect Anthropic (and OpenAI to a lesser extent) but I’m not going to play these games.
And a local AI solution is less capable and also quite expensive, but for some worth trying.
For 2-4k you can have opus 4.8 at home running faster than anthropic. In 8-20 months you have broke more than even.
I think for almost anyone it's worth trying. Especially if you already have hardware.
So trade-offs are there, but do those matter to all people the same way ( are they not equal in the same way to everyone )?
Because there are a large number of use cases (and large sub-portions of others) where genius-level AI isn't required.
Which means once that threshold is surpassed in people's relevant fields, available margin on that is going to collapse to commodity levels.
It's difficult to see how either pure-play AI company maintains its valuation once that happens. Their TAM is based on capturing a big chunk of all work, not just that which requires the highest intelligence.
All this indicates to me that loads of GPUs were only purchased on paper or are sitting in warehouses unused. This almost certainly means a drop in orders followed by a drop in RAM prices (though that would indicate a demand drop to investors, so maybe it’s better for stock prices to keep paying to bills bigger warehouses and stuff them full of unused GPUs).
The second RAM prices normalize, local LLM becomes much cheaper. A machine that doesn’t make much sense at $10-15k suddenly becomes a lot more feasible at $3-5k. If those stored GPUs flood the market, we might see even bigger price drops.
First, the form factor of GPUs in data centers aren't the same as desktop GPUs, so you couldn't use them even if you wanted to in a normal rig.
Second, from a business perspective that doesn't make a lot of sense. It's much more likely that GPUs are going to the highest bidder/large contracts who are scooping them up to populate/upgrade data centers that are in operation because they are going to get more money per gpu on newer hardware. The "old" hardware might go to a warehouse to be auctioned to the highest bidder or go to a data center coming online, but I highly doubt they are sitting on market wrecking amounts of GPUs just waiting to flood the market. When they go bankrupt and they have to sell a datacenter or two, those datacenters being parted out as part of a bankruptcy deal I could see. Warehouses full of unused GPUs doesnt make any sense to me though.
With eGPU PCIe 3.0 or 4.0 x4 links (or even Thunderbolt), that essentially doesn't matter for inference.
You pay the bandwidth hit on model load / unload (a few seconds), but it's irrelevant for post-loaded inference.
Anyone hosting a high wattage GPU is going to be fine hosting in an external enclosure, most of which have generous extra-spec room.
And if they don't, if a flood of cheaper DC GPUs hit the used market, you can bet Chinese manufacturers will have enclosures that fit them available the day after.
Yeah but it's unlikely to happen in coming 2-3 years (at least) and then the SOTA models are going to be a completely different animals.