When used with crush/opencode they are close to Claude performance.
Nothing that runs on a 4090 would compete but Deepseek on openrouter is still 25x cheaper than claude
Is it? Or only when you don’t factor in Claude cached context? I’ve consistently found it pointless to use open models because the price of the good ones is so close to cached context on Claude that I don’t need them.
Things get a lot more easier at lower quantisation, higher parameter space, and there's a lot of people's whose jobs for AI are "Extract sentiment from text" or "bin into one of these 5 categories" where that's probably fine.
And without specifying your quantization level it's hard to know what you mean by "not usable"
Anyway if you really wanted to try cheap distilled/quantized models locally you would be using used v100 Teslas and not 4 year old single chip gaming GPUs.