I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
At work, as an experiment, I used GPT 5.6 Luna + Deepseek 4 Flash for a week (I have an unlimited, "within reason", budget at work so normally I just use Fable and Sol) and it's been perfectly fine.
These models take a bit longer (more turns) to solve problems so they feel a bit slower but the end result is often just as good or nearly as good. Because they're so cheap you can easily run multiple sessions in parallel so it doesn't really matter that they're slower.
I've done a few experiments where I've split my terminal in 4, launched 4 clients (each with a different model, including Fable and GPT 5.6 Sol) and compared the output. For simple and medium complexity work open-weight models are incredible effective.
I can highly recommend the 10 USD/month OpenCode Go subscription. It offers pretty amazing value for the money and is a great way to experiment.
In a normal Chat with Fable, something like "How can I exfiltrate a guy from a sticky situation?" reliably downgrades, leading me to believe that Fable just outright refuses once it sees the word "exfiltrate". When it writes a service and names it CloudExfiltrator, the next turn downgrades to Opus.
Opus doesn't appear to refuse on simple keywords, but it does seem like Fable's reasoning introduces enough nefarious-sounding context that Opus will then refuse, and I'm stuck playing the new session game despite having done everything correctly myself and having a totally innocuous prompt. At one point, Opus was happy to continue while outputting commands for me to execute on its behalf, but flatly refused to execute them itself through multiple new sessions. To its credit, it openly acknowledged how ridiculous that was and was apologetic for the safeguard.
I'm open to the idea of some kind of guardrails, but if Fable is so dangerously intelligent as to require the guardrails you'd think they could come up with something a little more nuanced than a list of bad words. As far as I can tell, they've also not done anything towards improving the situation since the model was released, despite the "deliver more capabilities faster" claim.
Cybersecurity capability might be nerfed
(no personal opinions of either, links might be useful)
I think that OpenCode is nice, their CLI version is enjoyable and their desktop/web version is okay:
I also quite like driving OpenCode through something like Kepler / Paseo and tools like that (with those I can still use my Anthropic Condition by Claude Code being treated similarly - as something that gets tasks dispatched to it, while the GUI I see is Kepler / Paseo).
On the desktop side, ZCode was surprisingly usable for something that came out of nowhere (I wasn't aware of it at all before trying out the GLM Coding Plan): https://zcode.z.ai/en
(however, it's much faster to use Plan Mode to build a plan of what it will do, and then execute the plan in Build Mode. you can also have the AI make a script that will be executed deterministically)
When I install software on my computer with apt, I trust that all the files will go to the right place and install scripts are going to do sane things relative to the rest of the system. And I can just uninstall the whole thing with one command later if I so choose.
If I curlpipe a script, I get none of those guarantees. I have seen curlpipes that put files in weird places, guess the wrong OS, and mess with config files that I didn't want them to touch. When they break or I want to uninstall, I have to sit down and understand a (possibly minified) script to clean things up manually.
Yes containers are a half solution to this, no I don't want to use containers 100% of the time.
Gonna try to find a way to use this at work.
Like a VPN id kinda like to know? I’m happy to help fund a crowd source campaign for it.
I only found this yesterday, and it inspired me to start testing out OpenCode.
That said Claude Code is perfectly fine. I just prefer the integrated experience of using Cursors since I already use VSCode, but I still mostly use Claude Code because of their Max/Fable plan.
this is the integration branch for https://opencode.ai/v2 . it has been for months. it's where the Effect-based refactor has been landing.
CC works but for me it felt like increasingly they have zero incentive to make it a great experience. You hear folks like Boris talk about spinning up thousands of agents over night and agents chatting back and forth in GitHub issues and while I think it’s great from figuring out what the future looks like I don’t think it represents the reality of ROI today. So the folks building the tool are so disconnected I am simply not sure it’s a great experience anymore.
Is that the main concern though, cost?
That being said, I had to nope out of a similar thing from GPT 5.6 today, so it appears to be a US frontier lab issue. Claude is particularly bad though, as it produces far too much code even when I tell it not to, unlike GPT (and Kimi) which at least listen to me a little better.
More generally, I want a usable human review experience, and Claude code doesn't deliver that for me.
IMO part of it is that the underlying LLMs have gotten better enough that harnesses feel better even if they haven’t changed. I have a toy harness that barely implements the features you’d expect and it works surprisingly well. Like there’s literally nothing clever, it calls tools and that’s about it, and it still mostly does the right thing.
Edit: lol, I don't think ACP is even actively developed anymore. It seems to have been merged into another seemingly pointless standard with an even worse name, A2A. [0]
i like to challenge my assumptions and try new tools
that's a very compelling use case, thank you
It works nicely in the browsers on my tablet and phone, too.
On exe.dev you can ask it to customize itself, and it will automatically rebase your customizations when upgrading to a new release.
And this is coming from someone that's not particularly a big fan of Theo. T3 Code should get more recognition; people aren't just aware of it yet.
Anyways, please try mine!
I’ve stopped using it completely now.
What’s the counter argument? pi and ohmypi are pretty fantastic. Of course like all developer tools it depends how you do your work but I am not sure what you are trying to achieve in your comment.
https://stencil.so/blog/the-harness-problem
Three GLM models are mentioned, but so is Deepseek, Grok, Minimax, Kimi, and Gemini.
I think creating your own agent is the Hello World of agentic coding. Instead of Rust, I used D for mine.
Building [=======================> ] 610/611: dirge(bin)
Just hangs there :(
The tool calls will be, among other things, something like ReadFile, RipGrep, PatchFile, Shell.
When people talk about the value of different harnesses, they're also implicitly talking about the quality of the system prompt.
The same exact model, when given a different set of tools and a different system prompt, can behave differently.
The agent/harness is the sotware that leverage this "dumb" autocompletion engine to do useful things by sending the good input to the model and doing useful things with the output.
If it won't attack my stuff, it won't help me build my stuff to be secure.
What do you mean with this? Honest question!
I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.
In this case, you can put whatever you want between the harness you're running (or modify the harness itself), and essentially "lie" to the model. Any verification technique would be fairly trivial to bypass, while you continue to run the harness locally.
Have personally tested this with Opus and Sol and it works.
Classifiers are tricky though. Here's where open weights will win.
I've joined TAC, but still have to dance around it.
You can't guarantee everyone else will use a neutered model.
Anthropic was stingy as hell with its Fable and cybersecurity nonsense, switched to OpenAI which is much better but still not enough. I'm tempted to switch again...