Fable seems to be generally more impressive at outputting one-shot web apps. I'm not really saying that to try to downplay what Fable can do, it's just that if I compare the two, this is one of the few definitely noticeable areas that you can easily demonstrate. Obviously, one-shotting programs is much better as a demonstration of a model's capabilities than it is practically useful (not that it is useless, but hopefully my point is understood).
However, whatever Fable truly is better at, one thing I really like about GPT 5.6 Sol is even harder to quantify: taste. GPT 5.6 Sol outputs are still LLM outputs and they contain many things that people would probably consider "Claude-isms" for better or worse, but overall I really prefer the GPT 5.6 Sol output. I find it to be generally more tasteful. Hard to quantify, but when talking to people I've had enough people seemingly agree with me to convince me that it really is true.
I was very pleasantly surprised to find Sol wasn’t obstructive over what was clearly a very grey area endeavour.
Some of the topics it’s flagged have been hard for me to understand what it seeing that can be remotely concerning in my requests.
Which is definitely protecting their turf, but also probably a little bit hiding their “RSI” abilities for competitive reasons. My theory is that a lot of “safety blocking” is actually WIP training of new business directions. Anthropic has started hiring biologists and has opened a preview of a “Claude code for bioinformatics”. I’m guessing they’re tweaking their bioinformatics market play, and block “bio safety” requests so competitors can’t learn about their training.
That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.
Competition is competition and it’s two sides of the same coin.
Sol is great and has never blocked a request, and generally gives great answers. Happily switched over to it now.
There's people that have tried to contact Jeff McBride and follow the IP trail but the IP is currently owned by a company that went defunct. Not sold, but no one is even bothering to register its LLC any more, it's simply dead.
The only obvious objection I can think of to this line of thinking (at least from a technical perspective) is "how does someone build up the knowledge to be able to use a tool effectively in that way if not by doing things by hand at first?" The honest answer that is "I don't know, but that's also pretty much exactly the type of thing my employers have never been paying me to solve in the first place". Even just a decade into my career, there have already been plenty of times in my career I've struggle to convince people that we should do stuff in a way that won't bite us in the ass a month or two down the line, and in the times I've managed to succeed, it's usually only by putting in more of my own time and effort to make the initial investment seem more palatable. Luckily right now I'm not in one of those times when I'm having to go full throttle to keep the lights on a few months from now, but I don't have enough fuel in reserves to work on a plan for when we need to build a new rocket in another ten years. Maybe ask me next month.
MAI also offers a ultra cheap version that's competitive with Luna.
So much so that the models look like they were designed by a product manager explicitly to eat away OpenAI's market share.
Vscode even pushed them quite hard onto users with the latest release, going to the extent of putting up a modal to convince users to try them out.
Compared to what I am doing at home experimentally, I feel like day-to-day work is absolutely nothing. Not only am I also working with existing codebases in my experimental prototyping, but I am also doing things vastly more complex with vastly harder constraints.
But I wouldn't trust lower tier models for end to end solutions.
Sol over engineers now and then (hey please don't factor that function out into its own file....) but it doesn't do the same level of stupid terra does.
That said, plan with Sol, implement with terra, have Sol fix all the mistakes, then I go over the code and make recommendations for the architecture to fix Sol's foolishness.
Getting the AI to output code that you like is difficult.
As an example, let's say in React you have a "useLocale()" hook.
The AI will happily pass down locale as a prop to 5 child components instead of just calling the hook in the component.
A review from another model did not flag such stylistic issues either.
I believe that the latest models are very good at functionally achieving the goal, but still have poor taste for UX or code quality.
The most productive use of AI for software development happens in an environment where you do not review the code but test the UX end to end.
I think it sometimes worked, for example for testing preferences, but sometimes it did not.
Could be a problem with the harness also.
In any case, I feel that it's a bit playing whac-a-mole with explicit rules for things that a more intelligent model should do by default.
It's just vibes
It's... really just vibes?
Always has been.
When people figure out any reliable strategies to test and benchmark them, that's insane, and in the positive sense. This very same issue has been a thing for humans as well forever, and remains only very questionably solved (IQ, academic tests). This is not easy.
Just consider your own example. Do you think a less or more "rebellious look" is not something designers can actually ellicit? Less so in software design, sure, but in character design for example? Or general product design? Do you think e.g. Monster energy drinks are branded the way they are completely due to happenstance or something?
Except people don't usually put numbers to it, because they understand that that's hard to defend. You're the one who's describing such an idea, and wants such a thing to happen, classifying anything else as just vibes (that's the point!) and unhingedness. You're handwaving the difficulty and fundamentally limited nature of that, assuming that it is some laziness or mental delusion that's preventing it instead. You're also pretending as if it was somehow not real as a result. What I'm telling you is that you're wrong about that. Any kind of qualitative analysis that's actually defensible with these is genuinely difficult and limited in nature. See also all the opining about benchmaxxing. It also doesn't mean they're useless though, see also benchmarking.
The guy above didn't put numbers to his vibe assessment, they just drew a comparison, exactly because they know that there's not much else they can earnestly offer. You're sulking at them not lying to you by overstating their rigor, and you're flipping the arrow as if this limitation was some sort of mistake, not a necessary and intrinsic property, which it absolutely is. Natural language is an inherently subjective medium.
[0] In fancier and more mathematical terms: https://abeljansma.nl/2026/07/10/truth-is-not-a-direction.ht...
Where you’re wrong is pretending there is any intellectual rigor to the discussion which justifies promoting from the domain of vibes to actual reasoned debate
To give you a practical example, small, self-hostable models are very popular on HN. But my personal experience with them has been absolute dogwater, so this difference informs me that this is not the place where I should shop for a signal on whether a specific model like that is worth trying, or on whether it's game over for large models and remote models yet. I can "take the temperature" and make use of that without it having to be any rigorous, high assurance or mechanistic thing. It's suboptimal, but not useless.
Conversely, it also inspires people to try and substantiate these issues, so that it can eventually be more rigorous, higher assurance, and mechanistic. This is why I brought up benchmarks and benchmaxxing. Hard to know what points of consideration are salient when there is no abstract grievance to investigate, but Goodhart's law does also keep looming.
If people start moving away from Claude to GPT because e.g. Claude's output is too hard to work with and parse, that's relevant for the respective model providers, because it's a revenue share shift.
If I really like using Claude and strongly prefer its output stylistically, then claims otherwise will infuriate me, and will drive me to substantiate. If for no other reason, then because it's likely that Claude will have its language tuned in response to people's feedbacks and the revenue shifting, which may not be to my liking, and so I better prepare to call it and contend it.
It's literally like any other real world thing ever. Think tuning video codecs. Or tuning user journeys in frontend development.
But human coders can have bad taste too. There is code where there is nothing obviously objectively wrong, yet the choices feel like they were made by someone who just doesn't value or put emphasis on the right things, yet spends a lot of effort on trivialities. It comes in many forms.
To be honest, in this case, I wasn't even personifying them, because I in fact didn't say that an array of GPUs has "taste", or in fact even that model weights did. I was saying I preferred GPT 5.6 Sol's outputs as a matter of taste. I can see why someone would confuse the two statements since in this case they're pretty much the same thing, but if you re-read what I said I was actually more careful than you're giving me credit.
After extensively using both on Max 20x plans, I've concluded that Fable is better for problem solving and coding, whereas Sol 5.6 Ultra shines in debugging specific issues: tackle a problem with Fable then leverage Sol to clean up, double check, or fix specific issues.
Fable (imo) had the edge on the $200 plan, but after this 50% reduction I'd say Codex is better value by far and there's no contest.
---
Using Fable as the orchestrator and delegating tasks to Sol 5.6 Ultra via the codex plugin in Claude Code yielded good results, but still there was a lot more over-engineering (thus time and tokens spent) than Fable by itself would've done.
Both models suffer from doing-too-much. But both models are fundamentally really smart and knowledgeable. I think it's really close and pricing cuts really spice things up for us consumers! Sol is a clear winner in the value department and the $100 plan is enticing!
---
*Claude Code usage is reducing by 33% in 2 days, Wednesday August 19... cmon anthropic: clau.de/cc-50-promo
Given the 50% discount on Sol and how smart it is, yeah it's unprecedented value. If you only want to use low effort, there's a clear winner here on value and it's not even close!
Is it more about just avoiding any mistakes? Seems like that would be costly when medium or high would work fine?
For small tasks, you can just use something like low or medium effort and it can usually avoid mistakes; after all, the model will test the code anyways and can do some baseline level of iterating.
In regards to cost, we need to acknowledge how generous OpenAI was in the last couple months with Codex usage credits (no weekly limits) and usage resets. It afforded me many a dive with Codex! Yes it uses more tokens, but sometimes it's worth it -- just depends on what you're working on.
Finally, Ultra(code) isn't that bad when it comes to cached tokens. I think folks overstate the general token usage of ultra effort on both providers.
---
Both models are great at green-fielding a project when given detailed specs.
Both models overthink too liberally (imo) during these larger multi-shots. Sol overthinks more than Fable.
Both models are really smart and perform great for general knowledge and regular coding tasks.
However, I think these are very different models in terms of orchestration. Long-horizon tasks are way more predictable with Fable. It just doesn't lose track of details. Thus I ended up building a small wrapper around Pi (where I run Sol) so that CC can delegate via background tasks, automatically wait for completion, and do what was one of the most effective parts - steer Sol toward simplicity, getting Sol out of code-review infinite loops (Pi calls for Codex review to ship better, but generally gets stuck on P2 and results in vastly overengineered work).
One of the worst experiments was enforcing coverage at 100%. Only Sol, with an enormous amount of code and significant pushback (on architecture decisions) to Fable, was able to reach it. It made me think this is somehow related to overengineering in general, so that instructions on acceptance criteria in claude.md plus proper DX (e.g., Lefthook) actually led to okay results. It mostly helped that responsibilities were clearly split: Fable designs architecture, Sol handles coding and debugging.
Remove the one thing that differentiates these 2 SOTA models (long-horizon task scaling) and you're left with assessing the raw intelligence of both models.
Both models are really fucking smart. We should be intentional when discussing effort levels when it comes so SOTA models because currently, effort levels are the essential lever to evaluate task scaling.
I've never seen that leaderboard link but I think my sentiments reflect exactly the findings: Sol 5.6 is smart, Fable is smart, but Sol is more value for the end user (even more so when you lower effort levels because that doesn't degrade the model's raw intelligence/knowledge). Not to mention Codex resets, that's just the cherry on top!
But smart != capable and this is evident once you start assessing both models on long-horizon tasks with higher effort levels. While Fable is (imo) at least a little bit better, both are still very good. If you want to test the raw intelligence of said models, you should lower the effort level.. if you want to test the model's capabilities fully, you should increase the effort level.
Economics aside, Fable is the better model (imo), but there's no need for a binary stance here. Both models are very good yet there is a clear winner on the value front.
I'm in Australia, and Fable downgrades to Opus when testing for bugs in memory in a legacy C code base. If Fable starts taking initiative and writes a test case that involves writing to a null pointer, that's the end of the conversation.
What on earth are you asking it?
Also Opus 5 is fine if your codebase is simple.
> even after completing the verification program
Was it easy to complete it?
I ended up in some weird state where I can't even attempt the verification at all. Opened the Persona tab once, closed it and then it never opened ever again. It says a verification precheck failed.
Even without TAC, Sol doesn't seem to get blocked very often. Fable would downgrade to Opus if I looked at it wrong.
I hear you on the downgrades, I'm 13/13 on downgrades, and last downgraded me to Sonnet for asking for reasoning chain.
Its a crutch that is no longer competitive
It the first model to actually make me pissed off to use AI. I absolutely hate the model so much.
I don't even want to see the codebases this model is fucking up.
It might just be good at finding bugs that about it. That all I would ever use it for just because it works harder than Claude models.
I will not be told what I can and can't do by AI and I will no longer be supporting American companies run by despicable people. GPT only gets my money right now because its so fast and cheap but I'll be back to Chinese models in no time.
I just can't get Opus (Opus 5 is dumb as a rock, to be fair) or Sol to do that, so I exclusively use Fable for personal work. When I reach my weekly limit, usually on the last day close to the reset, I just go back to coding by hand ¯\_(ツ)_/¯
Heck, it even does security reviews and fixes, as long as I don't ask it to "attack" the codebase. I'm planning on using Kimi or GLM for that part.
They STILL don't have an option to "Sign in with Apple" on the website, but they do for Google??!? (and on iPhone of course)
Screw that asinine UX
(and no it wasn't better than Codex at this particular task)
Can somebody at Anthropic tag claude in slack or whatever goofy shit you do and ask it to add Apple OAuth to your website? Clearly humans aren't testing it.
and sure enough, I was right to do so: They don't even let you remove your payment method afterwards. Every other store, Steam etc., lets you.
No way I have enough trust to install their desktop app after that, so I just want to try it through their website..
Can Sign In with Google, but not with Apple
so you gotta open the Passwords app, copy your random email, paste into the website, then copy the OTP from your email..
It's been that way for at least a year
The desktop app was clunky too the last couple times I tried it a few months ago
and the AI itself hasn't been that hot compared to ChatGPT/Codex either: https://i.imgur.com/jYawPDY.png
So all the Claude hype posted on HN seems like a case of the emperor with no clothes to me
(P.S. The thing I just now tried to do on Claude hit the weekly usage limit after 2 minutes)
Sol also doesnt _really_ work but it sort of tricks me into thinking it does more convincingly :p.
cancelled my subscriptions few days ago. (was on 100$ ones, not sure if there is diff in quality for higher tiers or not.. there might be that too).
what i hate the most is that they will make any obvious mistake you do not tell them to avoid. then on the next plan to fix it, your token limit is hit at step 4/5 -_-. Both models seem incredibly good at that mostly...
for tasks outside of coding and program design i do find them quite useful. like devops crap. maybe because i hate that, i like their help there more.
I have witnessed 5.6 Sol Ultra edit line after line of literally empty lines ... for hours.
I wasn't literally watching it, I came back to a goal (that it started for itself without my approval!) that had done nothing but that for some reason.
It couldn't explain why it had started.