I've never had an issue with Codex or Claude reading massive files, they're really good at precise greps.
I've never had an issue with Codex or Claude reading massive files, they're really good at precise greps.
Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).
I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".
There’s an added benefit that the manager’s focus on strategy and task decomposition before actually handling the user’s prompted task directly seems to be a very good way to interact with Claude’s Fable safeguards, and I haven’t had any refusals doing this.
And while I haven’t ran any numbers, I can get orders of magnitude more out of my claude subscription doing this, especially with deepseek-v4-flash being as good as it is for as cheap as it is.
This is a Claudism, right? I feel like I never saw "gated" used this way before it.
Just a guy trying to make his subscription last longer than the single Fable prompt anthropic includes for 100 bucks a month, lol.
The community edition was open sourced when the creator got hired by OpenAI a few months ago.
very good way to put it.
That property would be very useful here, but I don't see how it would be achievable using LLMs.
So fable and opus use opus to explore. Sonnet uses sonnet.
I replaced my built in explore agent with one hardcoded to sonnet low effort.
Why not, though? I started using OpenCode + GitHub Copilot, but I burned through my Claude Sonnet quota in just three days. I switched to GPT-5.4-mini, which uses far fewer tokens, and it’s often just as good as Sonnet. I think optimizing token usage is a good exercise. We often assume a model will be terrible, when it really isn’t.
“Often” doesn’t sound great. If the smaller model fails then I just wasted a lot of time and tokens.
And why stop at 90%? I have this one weird trick to reduce Claude Code token use by 100%: use a different harness and model!
You kids don't get it, it's not about the tokens, it's about the principle of the thing. No self-respecting real programmer would accept the loss of even a few tokens over programmatic efficiency and cleverness.
The app would start using it for exploration tasks, and then as it improved it became the default for writing code and tests too. You can change it of course, but I find it does a pretty decent job if you have a large model directing it.
The parent model of course checks the work, but most of the time the handoff is good enough that no edits are needed.
It's also pretty fast and cheap, firing off a bunch of sub-agents to explore different parts of the codebase is a regular occurrence for the way I work.