(I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)
Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.
Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”
Once hyperscalers stop buying in the quantities they are now, there's going to be a lot of hardware supply to serve by then very hardware efficient models.
My day job doesn't have a dedicated enterprise contract with any of the ai vendors so I might trial it to see if it is worth promoting at the company. Part of me is still holding out hope that qwen 4 dials back the overthinking on its own.
Ah, I overlooked that, since I have a claude sub at work and all my explorations are purely personal. There's another fast version of 3.8: ThinkingCap[1] by bottlecapai but they have the same $1M restriction from what I can see since it's distributed under a PolyForm Small Business 1.0.0 + BottleCap personal-use grant.
[1]: https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.8-27B