I completely agree about SOTA, but it's a big leap from "you don't need Fable" to "you can get everything done with local Qwen". As always, it depends. Most LLM users are better off with a subscription (or even API pricing) because they won't use AI heavily enough for the hardware to pay off. Then there's the power users who benefit from larger models (software devs, for example). You can argue that there's a middle ground that would do just fine with local models, but I think this group is vanishingly small.
> The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
Optimal in what way? If I'm having to run tasks twice because the local model effed it up the first time and I'm resorting to my SOTA "backup", that's a waste of my time and far from optimal.