Plus I'm also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterday. Never seen anything like that.
The sharp but over eager jr. dev is a very good analogy :)
200L That's crazy considering the volume of a 1U server is what 15 litres or so?
But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there.
But simple example: if you ask Opus to do a review of the codebase (with a short prompt and not too much guidance), I've had it basically read the `git log` output, do a simple `ls` and have it declare "Everything looks great! No problems found!", when Fable really does what you would expect it to do.
And you might think: "oh, so it's just capable of handling crap prompts?", well sure. But even if you make THE PERFECT Opus plan (a plan that would take many turns/hours to finish), Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ...
If you give the same plan to Fable, it'll just DO IT. And it WILL get it done. And in the end it'll tell you "Oh, I also found 30 other bugs and I fixed all of them properly" (where Opus would have started crying, or WORSE, worked around the bugs)
Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work is getting done by very narrowly focused batches of subagents, with the main session just coordinating and keeping the doc updated.
Me: Hey Fable, I've got this massive, theoretically challenging, totally novel, ill-defined cutting-edge problem that I'd like you to solve.
Fable: < Doesn't merely solve the problem -- utterly obliterates it. Nukes it from orbit. Does a robust one-shot that takes several hours to complete. >
Me: Holy smokes, that was amazing!!!! But the formatting could use some simple refinements. Could you change the margins and maybe add a drop-cap at the start of each section in the user docs?
Fable: < Commences another multi-hour nuclear exchange with the code >
Me: WT?!?!
(The moral of this story is that bringing a nuke to a knife-fight is only occasionally the best strategy. And in more practical terms: Fable is amazing -- but only for certain classes of problems, and even if it were free there's a lot I probably wouldn't use it for.
Fable has been more intelligent, with better taste and defaults (e.g. make impossible states impossible without being told, build for testability), and considers/solves things that Opus did not.
My workflow is to run Claude in planning mode first to spit out a plan file and then review->revise cycle it with Codex or other agents.
One big tell is that Opus will say that it can't find any more revision advice for a plan file, yet Fable will find more issues but also smart pivots into better solutions. This is probably the best test since it's not based on vibes.
For example will inexperienced or experienced users see a bigger jump in subjective quality?
Less experienced people tend to use very broad prompts.
Experienced people tend to understand the structure of the code and give explicit guidance such that a larger model isn't necessary to read between the lines.
I noticed with GPT-5.6 (through work), I could step up my specificity by a level of abstraction. But I still intentionally scope the prompts fairly tightly, as I find it produces better results if you need to own and maintain the code.
In all cases, Fable clearly outperformed Opus.
I’m downgrading tomorrow.
It’s horrible slow and it feels like opus very often. It’s a totally different experience from the first week
Then I checked /usage and discovered I was still running Opus 4.8 xhigh.
Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.
Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)
given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.
unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.
But yeah opus often the better workhorse given price gap
1: tying up loose ends testing https://github.com/HarbourMasters/Shipwright/pull/5838 (fix: https://github.com/HarbourMasters/Shipwright/pull/5838/chang...)
I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).
Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.
Any chance you would elaborate?
Anecdotal, sample size of 1.
The only reason I tried fable was because Opus 4.8 went down the same line of reasoning about it as ChatGPT did. Fable solved it a lot faster than the other 2 spent looking into "false clues".
My weekly quota resets Sunday morning, so Saturday morning I upgraded to a 20x Max plan which also reset my quota. I burned an entire week of Fable credits on Saturday, my quota reset again, then I burned another week of Fable credits on Sunday. Both days were a mix of building features, reviewing code, fixing bugs, adding tests, etc, so a decent mix of real world usage.
The main takeaway for me is that while Fable is definitely a better model, the improvements from the model itself feel like maybe 10%, like this could have easily been Opus 5 or even 4.9 without all the marketing theater around Mythos and no one would have thought anything of it. The rest of the improvements came from harness/system prompt and effort level changes so that Fable uses significantly more tokens/effort/sub-agents at lower levels than Opus does (which of course is entirely controlled by Anthropic at the harness level and doesn't really have anything to do with the model itself).
In my estimation based on those 2 days of work (or two weeks of work depending on how you look at it), Fable Medium is somewhere above Opus Ultracode in token and sub-agent usage on any non-trivial task (Opus Ultracode uses workflows more than sub-agents, but it's a similar idea). Fable Medium will quickly spawn 6 agents in parallel, each quickly using 150-250k tokens, then will use 300-500k or more tokens in its own context. Fable High uses even more as it seems to default to 8 sub-agents instead of 6 and more tokens in its own context). I didn't dare try Extra, Max, or god forbid Ultracode as I didn't want to burn all my tokens on one prompt. Of course this is situational, it won't fan out so many for smaller tasks, but the whole point was testing larger tasks that I previously would have used Opus Extra/Max/Utracode on.
I really don't like how Anthropic is obfuscating their model performance by playing with effort levels. They did the same thing between Opus 4.5 and 4.8 to show a bigger performance gain for each point release than they really had (especially after 4.6 IIRC), so you can't even compare the same model apples to apples let alone a new model. Obviously they do it so they can market big improvements with new releases, but its pretty clear we're at the top of the S curve on model development at this point and are now brute forcing improvements via higher token usage (I mean Opus 4.5 came out almost a year ago, and the latest Opus and now Fable models are only marginally better while using way more tokens/cost...same on the OpenAI side with GPT 5 from what I can tell though I haven't used Codex much I have used the GPT model APIs a lot).
I also did an N=1 test with the same prompt doing a large non-trivial change to the codebase (migrating from Sqlite3 to Postgres) with both Fable Medium and Opus Ultracode, then had a new Fable session compare the two PRs...it decided Opus’s was much better! I can link a Gist with the review if anyone is interested, but I can't share the code as it's a private repo. I really figured Fable would bias to favor its own code, but I guess not. And Opus costed less (in tokens and subscription limits) and took roughly the same time (though you can’t really measure time since it depends entirely on how many GPUs Anthropic allocates at that moment which constantly fluctuates due to usage, plus Fable seemed to have been getting way more allocation than Opus during this test period as Opus was running unusually slow all weekend while Fable was ripping though tokens).
Also on a different long running review task using Fable High in Auto mode (exactly the kind of use case Anthropic promotes for Fable) where it fanned out a ton of sub-agents then collated and reviewed all of their fixes it completely lost the plot (while burning something like 20% of an entire week's Fable tokens in the process over like 1-2 hours). Its PR ended up having a broken Frontend test, it incorrectly thought it couldn't run the Playwright E2E tests (different from the Frontend CI) in the cloud environment due to a Docker dependency they explicitly don't have, and when attempting to get it to fix its issues it introduced new ones and overlooked others. The usual LLM failure case for long running tasks, no different from Opus or any other model. I had to have its PR re-reviewed in a new Fable Medium session to fix it up, which it did fairly easily (I'm sure Opus could have done just as well for much cheaper).
That test and that review session definitely reduced my FOMO a lot, on top of just my general experience with Fable Medium doing all kinds of tasks. They're clearly brute forcing like 90% of the perceived improvements in real world usage (and I'm sorry but 1-shotting toy examples where it seems to do much better than Opus is not real world usage).
Since most of the improvements basically just boil down to "every effort level is Ultracode, but much more expensive and possibly worse results"...I'm just going to use Opus on Ultracode for those types of tasks and keep using Opus's lower effort levels for smaller tasks. Once they eventually add Fable back to subscription plans I might use it sometimes, but from my experience this weekend the improvements are absolutely not in line with the cost increase and I'm not willing to burn a whole week's tokens in a day just to use it when I can use Opus all week without hitting my limit.
Oh and one interesting observation, I never got kicked back to Opus by the security guardrails as far as I know (a friend who was getting kicked out a lot confirmed they do inform you and I never had that happen). I was even doing a lot of reviews for code correctness and bug fixes which I thought might trigger the protections, but never did, though I never explicitly prompted it to look for security issues or vulns.
Haha I just gave the exact same prompt to Opus Ultracode and it thought Fable’s was better.
Obviously this isn’t the most scientific test due to LLM non determinism, and I still need to manually review both to make my own decision, but the fact at least they seem to basically be a wash is pretty telling about how much of an improvement Fable is when you actually compare them as close to apples to apples as possible (aka similar actual effort/token spend/sub agent activity)
> atomic.chat (@atomic_chat_hq, 2026-07-02):
> Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8!
> We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
> Prompts:
> — A train derailing off a broken bridge into the water
> — Two cars jumping off ramps and colliding mid-air over a canyon
> — A monster truck crushing a row of parked cars
> Outputs:
> Fable 5: 62,158 tokens, $3.12
> GPT 5.5: 37,753 tokens, $1.14
> Opus 4.8: 22,280 tokens, $0.56
> GLM 5.2: 36,246 tokens, $0.08
> Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.
- https://xcancel.com/atomic_chat_hq/status/207244606796297841...
- https://nitter.net/atomic_chat_hq/status/2072446067962978411
There are more public Nitter instances at https://status.d420.de/.