Good thing then that you're fully in control of your own actions.
A gift granted to us by being a fully grown adult that is also likely registered to vote.
Meta: There is of course a way to put this less snarky, but that doesn't slap people in the face as hard as they need to be slapped in the face to maybe one day start remembering that they have agency.
I cancelled my Anthropic subscription until they fix how their models write and it's no longer unbearably annoying and obnoxious. The concise output format is a step in the right direction but I need a few months away from them.
Kimi and GLM models on Max reasoning feel pretty close to Fable. That said, even at their most expensive plans, a single one of them might not always be enough, while getting both of them for a year gives you a nice discount and isn't insanely more expensive than Anthropic. The problem there is that they could still easily rugpull you with token limit changes later, I don't trust any of the big labs not to mess around with those for any length of time.
Also most harnesses let you choose models per sub-agent. Like I can use Fable for running the main session and just tell it to use Opus agents for implementation in Claude Code, same with the Kimi and GLM models inside of OpenCode and other harnesses. The only problem is that the UI for controlling sub-agents usually really sucks.
Also, OMP has a solid subagent model and I like it with some tweaking.
(I'm in the second boat so long as I'm responsible for the code I PR)
Is there any alternative model with design sensibilities?
If you're happy to "lead it by the nose" you can do well with a lot of very low end models.
If you want to kick off a "/goal run until [complex verification passes]" and let it run for a week with minimal intervention, then not so much.
You can make do with cheaper models for long running agentic runs too, but it tends to require a lot of extra scaffolding and additional review steps.
For me, and my projects, it’s been great. It’s made an enormous difference.
I guess my workflow may seem “quaint,” to many folks, here, but the end results speak for themselves.
I suspect that one vocation that could get heavily impacted by AI, is the consulting business. That’s where many experienced people go, as they reach their career peak.
In my last project (just about to ship), ChatGPT replaced a whole bunch of services that would usually be supplied by external advisors.
But these are also services that I would normally not be able to afford, otherwise, and would just have to “make do” with. This release will have a level of polish that I have would never been able to achieve, unassisted by AI (I had originally used “on my own,” there, but the reality is, it actually was “on my own”).
But in my case, it wasn't nearly so exotic. The LLM helped me to do a much better job, preparing the App Store presentation, Web support, privacy policies, budget prognostication, and app glossary.
I have just had an extremely complex app, pass App Review, in record time (from going into review, to approval). No niggles or bounces at all.
I'm thrilled.
Basically, my needs are different from others. I'm not working on the next NORAD upgrade, much of my work is open, and the more ChatGPT knows about me, and the app I'm designing, the better. One reason I chose it, was because of this "memory."
TL;DR: I feed it just about every scrap of information about my project as I can. Source files, documentation, screenshots, videos, information about the organization, information about the target demographic, etc.
With all that information, it gives me very useful advice.
It created a great tutorial. I usually write way too complicated ones. It did much better.
In the case of the App Store stuff, it helped me to choose the right privacy report, generated the privacy manifest, and helped me to compose all the copy on the storefront.
I'll probably be releasing the new app, soon. It's already passed review, but I want to make sure that everything is kosher, before releasing. It came together so quickly, that I have the luxury of time. I just need to release before (or as) iOS27 comes out.
Even though I use MD/skills and other methods, I have remind models that they are not paid by the word multiple times a day.
For example, I have made a tutorial, which is meant to be a "quick reference," from within the app (Use Safari to view the page), but I am also developing a "walkthrough," to show possible funders (we're an NPO). The walkthrough is a higher-level vocabulary than the tutorial. The LLM deals with stuff like making sure to keep the glossary consistent, etc., but I like to have the final say on the output.
I'm pretty sure that I can force the LLM to use certain levels of vocabulary, through the .md file that describes the default setup, but choose not to do it.
I am still in that "trust, but verify" stage of my relationship with LLMs.
I don't have Fable at work but I'd probably use it for actual code if I did because not having to spend time handholding the model on this stuff and getting useful code first try is very useful
And Opus5 aggressive audits.
Once it has exactly your coding conventions and access to other code to copy bespoke patterns, a strong idea for what to do, then you can let it do the work.
You do the wiring, it fills it in.
Coding was never the work.
Beyond that, I find this whole plan and build thing to be a pointless waste of tokens. If your planner made a detailed enough plan, then the cost of executing that plan is a just one turn more of cached tokens, and minimal time.
Meanwhile: switching agents, reloading context and building from the plan will easily balloon your token use and time. And any emergent problem that the dumb executor finds will instantly wreck the implementation because they're not competent at solving it. And if your plan is so perfect that there's no edge case then you're wasting tokens because your planner was one turn away from finishing the project via cached tokens.
"then the cost of executing that plan is a just one turn more of cached tokens, and minimal time."
This is just not true at all. There's a huge gap between 'figured out the hard stuff' and 'rock solid'.
Dependencies, integration, corner cases, docs, testing, unforeseen issues, a lot of back and forth auditing making sure things are really tight.
Audits get diminishing marginal returns, but you have to do them until they don't find anything, and that's usually a few cycles.
So aside from the fact there is 'a lot of labour' - part of the plan (maybe the most important part) is documenting most of the trip-up scenarios. If you ran an experiment or two in the background your agent will 'discover' a few key odd things, you back those into the plan.
I'm 100% certain that this pattern works because I (and others) use it very successfully.
Hint: save your main context by using sub-agents to do grunt work - even in impl phase - farm out anything directly implementable without a ton of background.
Also - make a skill so your Claude can call Codex and visa versa and maintain long-running sub agents of 'the other kind'.
An Opus with 1M context window executing on a 'plan' that a Codex 'sub-agent' is executing on - ad a different Opus sug-agent is auditing hard ... that 1M token window is dramatically extended to 'many millions of tokens'.
That can work within Anthropic/Codex Pro plans.
I don't work where we can ship slop. I don't work where PRs can be merged based on what the agents say. I work where a human has to read and approve and own every single line of code. I work where the stakes are actually high, so the cost of not using the best tools in terms of human time are big. A single turn around in a PR costs more in human time than the difference between deepseek and fable in API costs.
So, when you admit "There's a huge gap between 'figured out the hard stuff' and 'rock solid'." but then claim that the cheapest/dumbest agent in your arsenal is your go-to for "rock solid", I have to question the quality of your results.
Personally, "using plan mode" is a very 2025 way of using these tools, and I wouldn't be surprised to see "plan mode" be removed from codex/claude code/et al.
Realistically, I'm using the best models to think about a domain and problem (Fable High+), and I'm using a cheap daily driver with an advisor pattern (Opus High + Fable) to iterate through POCs, and I'm using human review to guide design. None of that is "plan mode", it's actual engineering. Then we decompose the solution, we stack it, and we use only really strong agents to build, review and refine.
This obsession with cheap agents leads to low quality outcomes. "Rock solid" deserves the best tools, and the "plan" will never be good enough. I'm going to be sending fable xhigh and sol 56 xhigh et al at it in adversarial review, why the heck am I cheaping out on the actual implementation?
And finally: my time costs way more than any of this. Cheaper models are slower overall and when combined with re-work time, are dramatically slower. I'm costing my company hundreds in my time to save a few bucks on the API bills. Nonsense!
Based your arbitrary dismissal and unwillingness to even try to consider new patterns with which you may be unfamiliar - it may be difficult to communicate with you.
I have the advantage of 'certainty' because I have the evidence over many projects / team members.
We ship near perfect code.
In addition to the hints above, we do this at least in part by explicitly anchoring and testing requirements into several aspects of the code, and ensuring that known 'weak spots' are managed.
The 'planning process' ensures the requirements are mechanically anchored and integrated into tests, that 'proportional' documentation is applied, and that module, library and project level documentation is perfect (and mechanically validated where possible), which FYI is what solves most of 'context problems'. (That's another hint, if you have extremely good docs, you don't need to load vast amounts of code).
Yes - I hear you that 'time matters' and that 'the stakes are high' - consider that you may be talking to people where the stakes are just as high, or higher - but more specifically, this is not about 'saving tokens' or cost so much as it is using the right level of model for the task.
Use the best models for background research and planning, use mediocre models for execution, and mid-high for auditing - in other words 'use the right model for the right work' - and in a certain methodology, dumber models are appropriate.
FYI this saves you the ugly 'Fable' problem which many are encountering as it burns though Max plans. Don't 'automate' with Fable, it's the wrong model for that.
I could go on, but consider that there are actually ways of organizing projects and orchestration that work well.
I feel the value I get far exceeds $600/m. It's a straight expected value calculation for me and I'm happy to pay. I wish they'd make it easier though - just sell me a 100x account for $1k/m and save the messing around.
I tend to work in 5-10 or so parallel streams at a time so that also multiplies token use as a function of time. Much less and I'm waiting for it too much, much more and I can't juggle effectively. My new big project is of course to try to take myself out of the equation further and increase the parallel streams dramatically; pretty hard to get that right so still manual for now.
This is my current workflow: https://x.com/sridca/status/2093354491965735091
I even let Fable choose the appropriate model for the task.