Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model)
Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model)
The fact that it basically broke open the local model supremacy was just a nice side effect.
I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context.
Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need.
If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans.
What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier?
Frontier has to displace labour one for one at some point or it’s over.
I wonder if the next DS models will also graduate to 5.x, I think I saw they are training up a 10T model, and just raised $12B too
It's important to realise that Deepseek are a very unconventional model lab, as they were originally funded by the founder's work in quantitative trading.
They mostly don't need that much money, the reason they took funding is to provide employees with equity, as lots of their good people were getting poached by other Chinese model startups.
(I have no inside info, just read the FT articles about them).
In practice, you will be able to run models a bit bigger than 35B.
38GB of vram resident. more tk/s, more prefill.