That’s been fine for smaller tasks like analyzing a document, but for more involved work like refactoring code the latency makes it harder to iterate.
(the main reason is not just cost, it's data not going out and even mess leaving the EU, this essentially frees us of a lot of hurdles)
This is clearly pushing it memory wise and the gpu offloading is only partial but LM studio deal with it automatically and it's being fine even on the 4071 Ti desks (Ryzen 7500F and 32 GB of ram), employees get any feedback in ~20 minutes after they dropped a file (it's much smoother on the 5080 desks obivously), for live it's useless but as background helper it's great and the very large context allows us to fit all the rules we want in there.
One caveat has been to not ask it if everything is ok, but to find what's wrong - but always source and explain it and justify itself, never drown the user in warning in suggestions; goal is to help and provide a second pair of eyes not make them feel annoyed or unsecure.
And I found people to genuinely enjoy something that works for them on their own machine and is not tracked "by the boss", thus the assistant reference, than than a centralized mothership like we also have and they have access to. I also allow them to disable it if they want, I trust them with their work, but a second pair of eyes is always great.
(my previous workhorse for this was Qwen3-14b but it's missing a lot more edge cases)
Also, the system prompt in the article is supposedly for the web version, so I think the $$$ API version or providers that still allow third-party harnesses like OpenAI should have fewer limitations.
Better tech doesn't matter if it is non-tenable for the general public. Eventually, the higher volume product will win.
Eg if you use Claude, you probably want fable architecting, a couple of opus under it managing sub project and sonnet doing the actual function code, because fable coding a "run a query and filter the result" is a massive waste of abilities. But their own sub agent downgrade is limited to one level so if you use fable it will never direct sonnet coders.
I've been doing this myself, all the time. (The ability to get second opinions and code reviews from Sol and other models is priceless.)
There's nothing restricting it to only controlling subagents that are other Claude models like Opus and Sonnet.
Also, I'm not sure that there's a one-level downgrade. Subagents can be pinned to any model: