Anthropic's best AI model struggles to attract users as cheaper tools thrive
ft.com
ft.com
They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
"You can only use Fable for a week as a part of your plan" "Be ready! You have to start paying per token!" "Nevermind! we extended it for a couple more weeks" "Wait, now it's up to half your usage" "Ok, now its..."
Most people want to not care. We want our AI like electricity -- Kind of just there no matter how easy/hard is for the supply. You don't want your electricity company to be on the brink of cutting you off any second.
That's Anthropic. You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game. That forces people to look beyond the walled garden. There, they find models that are fine... and without the shenanigans.
It also didn't help that the government yanked it which adds another source of anxiety since OpenAI is on much better terms with the administration and the administration seems corrupt enough that they would mess with Anthropic if they got a big enough donation from OpenAI.
But anyway after Sol entered the picture, I don't think Anthropic can get away with this as much and I also think they're going to face a massive backlash from Max subscribers if they do end up ending the +50% promotion at the end of the month because Sol is a Fable peer and priced very competitively.
The thing I miss most about programming is flow, and the constant bouncing between terminal tabs sucks. I’d love to do one thing at a time, with Fable, quickly.
An ide open with 20 tabs open each file a component, a class or an interface We also use to hold entire codebases in our brain.
I'm struggling to scale myself even further. This tech is unreal and I have so many things I can do.
For the first time, tech feels like the 90's-00's again. Everything is greenfield and exciting and big tech is struggling to figure out what to do about it.
People are just hacking all kinds of stuff, and it's awesome. Feels like techno utopia.
The barrier and time between idea and usable implementation is almost zero now. I don't have to imagine. I can just write something and see it work before making larger decisions. I really like this. Many of my ideas were abandoned because I needed to study some obscure library. Now, I can learn the parts that I find interesting and just have the AI chew through the grunt parts easily. That's the good.
I started coding with a line editor on a small Casio handheld "computer" and used to keep programs in my head. I more or less knew what happened on each line without seeing the line. With larger programs, I had a mental model of what was going on where and a big part of the input to that was the effort of writing everything by hand. That's gone. It's not really important as far as the output of usable programs is concerned but there's a certain feeling of satisfaction that came with digesting a larger codebase and having it surrender it's secrets to you that's missing.
It feels like the opposite of 90s - 00s: they were filled with periods where a person could self-study technology and get a job using those skills that few others had.
Where we are going (according to the AI-proponents) is children being able to replace you.
In brief; the 90s - 00s were a skill-valuation time, now we are looking at a skill devaluation time.
Unless you meant to say "Just like how any kid who could write broken HTML t put up a webpage could pretend to be a skilled professional, that's where we are now"...
Then they hit you with hourly and weekly limits, you are wondering when you are going to get cut off. Since LLM at probabilistic it often feels like pulling the lever of a slot machine, hoping our prompt is the jackpot. To make sure we win, we come up with systems, convoluted agents, context pipelines, rags to load . It feels like it's working, then bam you reach weekly limits.
(Just pay more if you want to keep winning).
I think developers need to wake up.
I myself started to use AI like a fancy debugger ,explainer. I make it walk me though every single line of code it writes.
I notice that i run into limits less, if i get cutoff, i can still make changes
But for instance I used to work in 1-2 client projects at a time and they take months now I can do 4-5 at once and they take a month. That’s a huge improvement
I don't dispute that your tasks are taking less time, but I suspect that to be temporary. This is a forum of AI frontrunners and early adopters, so I expect it's a matter until people catch on, and recalibrate their expectations for amount of output a programmer can produce in a given timeframe.
What I dispute is that this increased output amounts to actual value. There may very well be a HN-wide 10x productivity boost, if HN measures productivity in Jira tickets per day. But if that's the only measurable result, we should expect the only long term change to be a 10x increase in Jira tickets once orgs catch on.
My skepticism is we're not getting that because the world mostly doesn't need more software and it doesn't need all software to be ultra-personalized to each user, either. There are plenty of valuable problems that may be solved via computation, but thus far, it seems we're getting the software equivalent of movie theaters disappearing and being replaced with people watching TikTok from bed. It re-routes the monetary value extraction from consumers of audio-visual entertainment, and TikTok probably has at bare minimum 1000x the content-length of the Criterion Collection, but it isn't making the world 10x better for anyone but the owners of TikTok.
It reminds me of my best friend from college, who was bipolar. His goal in life was to become a writer and he eventually did become an Emmy winner, but back in school, he's go into manic episodes in which he'd stay up all night five nights in a rows and churn out thousands upon thousands of pages of free-form text that incorporate prose, poetry, play scripts. It was definitely more than 10x the output of a non-manic period, and there were nuggets here and there of intensely evocative single phrases, snippets of dialogue that looked like they could come from a more compelling story, but they were ultimately sketches and drafts, not anything publishable that another person would want to read.
Would we call him 10x as productive when he was writing more total output as measured by number of words or when we was writing much less but in a form that millions of other people enjoyed and remembered?
If we purely mean economic productivity, that is pretty straightforward in a case like yours. If you were previously completing 4 projects every 3 months and now you're completing 42 every 3 months, are you earning 14x as much money as you used to?
heh this is funny because this reaction was all the rage in the 90s early 00s too. You'd put together something you thought was cool and then post a link on a forum only to be told how it wasn't even "remotely impressive". I'm glad people didn't give up back then and i hope no one gives up now.
Things just moved up on the abstraction-ladder, but it's still there, hidden beneath all the TUI sessions instead.
Even a year ago when i was trying to do a hugo template manually with the help of the documentatin /tutorial, it was shit. The LLM at that time, was better helping me than the documentation.
I had some helpful people helping me on IRC / Quakenet.
But the hugo example i found very interesting because it was the latest hugo ducumentation and I don't think I was able to find a tutorial. I tried it without an LLM first.
Then she hired a VA in the Philippines. Anthropic promptly banned her account without warning once the VA connected to the account. It took her weeks to get her account reinstated, at which point she had already moved on to OpenAI.
How can small companies with 1000x less money able to provide live support, but if you pay 20, 100, 200 dollars for a subscription you dont have a phone number to call ?
You message support, some real person reads and gets back to you within a day.
What your your company do? Is it low ticket business or a high ticket business?
Our free tier is time-limited but we still look at all tickets even from non-paying customers (in the hopes of converting them). A 1-hour intervention from a customer rep can result in a multi-year paying customer.
If firms had a base degree of customer support they were expected to provide, they would still exist. They would just not be as profitable, but customers would be better off.
I seem to remember there was a time when S/W was also designed with the aim to be easy to use, so that the need for support was reduced. It feels like the lesson learned was to keep costs low, not to ensure users were ok.
you'd end up with a call center larger than most cities. It's not feasible.
Sorry, you can't buy an iPhone because Apple has got too many customers.
Sorry, you can't have a GMail account because Google has too many customers.
If you could be profitable with 50 customers providing excellent support, you can be MORE profitable spreading that excellent support across a larger customer base.
Just because none of the big corpos choose to do this does not mean it isn't possible. It's JUST greed.
I can't tell you how many times I've experienced this with comcast. The last time I had to deal with it, was when I bought a new cable modem. I call in to provision it, the automated system assumes I have one of their modems and fails. For some reason I can't get technical support on the line and finally I resort to yelling 'cancel my account' over and over again until I finally get someone on the phone.
The guy was able to solve the issue in 5 minutes flat. The problem with automation is it's only ever going to be able to handle the 'happy path'
The solution here is removing corporate monopolies and political power.
Right now we're using OpenAI's models by default since (unlike Anthropic) will allow us to use our pro subscription rather than token metering, but I've already had the joy of being able to change it to Kimi K3 (via OpenRouter) for an hour to try it out, and there were zero hiccups.
The period was also marked with many billing bugs, like spending people’s usage credits for included Fable for a few hours (gave me a huge shock), but to their credit they refunded it.
Their disrespect for their users is also another problem. You only get one shot to make a good impression.
Same, it's been a while since I logged into Claude web.
ChatGPT web usage being separate from Codex usage limit is a nice touch unlike Claude.
Most coders don't pay for tokens themselves. It's just on reddit and HN you would think that everybody does.
(And failing that, there was a real risk for OpenAI to be the default for enterprises.)
To defeat that, you need to frontline employees the chance to experience better tooling and models which is where the subsidized subscriptions come in.
Work can pay a Claude sub. At home i see no need
I'm strongly considering biting the bullet and just ditching my $200/month Claude Code plan for the Codex one instead, especially because I keep running into my weekly limits (even sticking to Opus.)
1. No 5 hour usage limit
2. Weekly usage gets reset CONSTANTLY. It's crazy. The longest I've ever seen it go without a reset is maybe 5 days?
3. I don't feel like OpenAI is constantly trying to fuck with me. Unlike Anthropic. I would way rather have Sol all day every data, consistently, than a slightly better Fable for like, 1 prompt every 5 hours, and only when Anthropic decides to not treat me like a cyber criminal. Believe in yourself as much as Claude believes your CRUD app is going to hack the pentagon.
4. Getting access to image generation, though I don't use it too much, is a nice perk compared to Anthropic.
edit: Should mention that I had like 4 banked manual resets as well. It feels like OpenAI wants me to use their product, whereas Anthropic wants my money while giving me a nerfed experience
I have used Fable heavily on the lower Max plan and you are really exaggerating the limits here. I've had many multi-hour sessions with Fable on Max.
We all know that the subscription prices are not at all sustainable for these providers. You all do, right?
Yes, they're struggling to segment the market and find a way to make money, and that basically relies upon emptying the pockets of whales. As someone enjoying a hilariously subsidized Max plan, I understand that, and I don't think they're trying to scam me in some way.
And both sides of this equation understand that the market is competitive, and maybe more competitive than they thought it would be. Like, would you rather they did pull Fable when they first said they would? Or that they'd cut quota? I wouldn't. But I'm glad that Kimi K3 and GPT 5.6 Sol and the latest GLM and Qwen and...I love that this has forced Anthropic to change plans. I'm not going to hold that against them.
Yes, they are most certainly subsidizing the subscription plans. I mean, at least for people who utilize them to the quota.
I mean, this argument would have merit if they built something that they're going to monetize for decades, but cutting edge models now grow obsolete in a 6 month time window.
There is zero financial analysis where they aren't massively subsidizing the subscriptions, and people have to invent ignorant "oh just ignore most of the cost of providing the product" to try to pretend there is.
But the agentic layer is being worked on, agents will start consuming more and more tokens
I’m at the point where I need stability and predictability. I want the B- student who shows up everyday rather than the A+ student that’s unreliable.
Their truth is “we don't have compute and are working to improve capacity”
People would root for that
Instead they got people rushing to escape the permanent underclass until they have a mental health crisis just to beat the fake deadline. $100, $200, is a lot for those people
The consumer side cheap monthly plans exist for the same reason companies like Cloudflare and Vercel have a free tier: When it’s cheap and easy to get developers familiar with the tools, they will push their companies to pay the real money for those tools.
It’s a hard balance with LLM serving because you can’t really make it free. $20/month is close to free, but the $200/month plans are in a difficult place where they’re big enough that many small companies pay for $200/month plans for their employees and ignore the enterprise features you get with the full expensive arrangements. So the companies are continually adjusting the $20-$200 plans to keep them from being reliable options for businesses, which is where the real money is.
There’s a short sighted cheering on of the 3rd tier and lower companies offering lower rates, but we’re already seeing them ratchet up the pricing and keep larger models closed after they get market attention.
Sorry, but people are cheering on Chinese companies (of whom your are unduly dismissive with your '3rd rate' comment given how good GLM-5.3, Kimi K3 are) not only because they are more economical, but also because they do not constantly refuse to do legitimate tasks and provide you with the weights for self hosting these models.
Maybe Anthropic's enterprise sales are going brilliantly, and the rest of us are just pixel dust to them.
Still. Brand perception is a thing, and between rug-pull usage policies, weirding verbedly output quality, and "I'm sorry Dave I can't do that" pushback, Anthropic are clearly having strategy issues.
Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again.
What business encourages users to try their competition and adapt their usage to the competing products?
I guess this is why they're pushing Claude code hard (not supporting agents.md, not allowing third party harnesses, etc) but when switching to another provider is as easy as opening a new terminal and typing omp/pi/codex your moat is effectively zero.
They can compete on price, quality or value but anything else is just madness. Currently they (arguably) own quality but this won't last.
I've never run a local LLM model before. Certainly won't take as long to iterate on this.
I have only 48gb of ram, so can fit only 80k context max, so good compaction is must.
I suspect the same will/is happening with AI. Either you will pay for it, or it will be so ad infested that it will become useless.
https://support.claude.com/en/articles/15036540-use-the-clau...
Update June 15: We're pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed: Claude Agent SDK, claude -p, and third-party app usage still draw from your subscription's usage limits. The previously announced monthly credit, which would have been available to eligible claimants in connection with these changes, isn't available. We’re working to update the plan to better support how users build with Claude subscriptions. When we have an update, we'll share it before anything takes effect.
They are high on their own supply. The people running these companies are delusional imbeciles who have been placed in charge of billions of dollars.
If you are selling something, and losing $10 on each sale, you also would want to limit how much you sell.
I mean, sure, you are losing money on each sale so you can landgrab, but you still have to balance the land-grabbing with how much money you can actually lose.
From a cost-effective perspective the GPT models are much cheaper - one can easily tell it takes longer with GPT5.x to exhaust limits and this matters A LOT.
I can’t say which of these corpos I despise more though. I though for a while Dario was cool, but a massive distrust is piling and the first third player offering decent experience (and showing some decency) will win me over.
For the record - I’m also unsure whether I despise more Exxon or BP or burning fuel as a whole. Hope u get the point...
I just put $30 on openrouter, switched to Pi, and I finally have a calm mind. Since I actually pay per request I want to maximize efficiency rather than utilization
To me, all of these are just exercises in getting me to pay for more tokens at API rates.
It doesn't help that all their models are bow trained to waste as many tokens as possible with their extremely verbose output
Instead often it feels like they make a change, then wait for someone to figure it out. Then Anthropic ends up being reactive as opposed to proactive in communication.
It is kind of funny because surprises from OpenAI tends to be positive (Tibo resets), on the Anthropc side I dread them.
With newer models text generation outputs have gone from probably human readable text to dense philosophical treatise about "load bearing" and incomplete sentences. So much so that now you need skills or another LLM to just parse the output. Simple answers and text generation just doesn't exist.
The breaking point for me was the privacy violation. They've been fingerprinting every request and violating users' privacy hoping no one would notice. Too bad, someone found out and that was the day when I cancelled my subscription.
Its in my opinion not wide spread at all and as today is literally the first time i ever hear anybody mention this.
Anthropic have been anything but. Flip flopping on model availability, model access behind an opaque filter, their past behaviour of model degradation as they prepared their next model… these are not signs of a reliable service.
I’ve mostly settled on using a mixture of open weights models through Together.ai and Fireworks.ai, a MiniMax subscription for high-token-use tasks that don’t need the best model (for $20 I get what feels like infinite tokens), and codex for the occasional high complexity task, although with Kimi K3 and hopefully soon GLM 5.3, it’s becoming increasingly less important. Deepseek 4 flash is my cheap main with delegation to other models as needed.
I’ve also found LFM2.5 8B surprisingly useful for single-focus tasks like “does this diff touch anything that isn’t related to the task”, and it’s incredibly cheap ($0.03/0.12 per M in/out).
After Fable 5 launched, it was better than Opus 4.8 for sure. Then they rug-pulled Fable from me (EU), and later released Opus 5. Now I only reach for fable when Opus 5 API returns 529 for the millionth time this year.
This is why I think open weight models will win out in the end. Right now there’s too much going on behind the scenes with the models. Day to day you never know if you’re going to get smart Claude or dumb Claude.
I'm not saying this with any undue derision, it's genuine - do they have a real product team or are they clauding that too? The direction makes little sense.
The only B2C is going to be watered down ad-driven BS, and they will charge B2B via tokens.
The $50/mo - $200/mo consumer LLM subscription is not something I expect to last long / or to drive much of the revenue share... like individuals paying for Gmail vs Googles overall business.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
I just kept a $20 plan going for use on my phone.
I don’t doubt people are hitting it… shrugs
That would be absolutely insane... I don't think it's doing that?
https://readysolutions.ai/blog/2026-06-10-claude-fable-5-sil...
> "The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives."
From: https://www.anthropic.com/news/claude-fable-5-mythos-5
> "Some attacks only become visible across multiple requests. Best-of-N jailbreaking, for example, sends hundreds of slight variations of a prompt in the hope that one will work. Larger patterns of misuse, such as state-sponsored espionage or data extortion campaigns, only surface when our safeguards classifiers can zoom out across many requests. Detecting these threats requires temporarily retaining prompts and outputs so they can be analyzed together, rather than one at a time."
From: https://support.claude.com/en/articles/15425996-data-retenti...
(yes, it was a very lazy prompt I could easily have googled, but that makes the refusal even more bewildering)
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
Yep, all they've got to train on are bot-generated Instagram posts. Can't say I pity them, though.
This is the part that is truly scary.
Nadella the CEO of MSFT wrote that post “ A frontier without an ecosystem is not stable”
It seems to me that the unspoken assumption in that post is that no matter what happens they’re gonna be training on your data.
He’s the CEO of Microsoft, he knows how these decisions go down, he knows how the world works, he is sending a warning.
The frontier labs can have the models but without being where the workers are they cannot do much, lots of industries have strict requirements of not sending their data over to randos in the internet.
Microsoft and Google have the upper hand here with their workspace offerings and could easily position themselves as secure enclaves where you can use local LLMs where your data never leave your premises and is never used for training.
It's not even that expensive because even the 15€/month/employee bill for people who use AI very little can stack up.
Beside that though, many companies already have most if not all of their data in some cloud (Microsoft would be a prime example). Giving them a few extra bugs to get a (potentially) very useful tool is not out of the ordinary.
If you are under the impression that not going AI would threaten your business right now (which may very well be true for some companies) then there isn't really a choice even if you believe the vendors will steal everything eventually (which they totally will)
Complete nothingburger. Actually worse than that, its a London Horse Manure crisis.
>Destroying millions of obscure books to scan them?
I really don't see the issue. They got slapped in the face for trying to do things the right way and torrent the lot. Why wouldn't they exercise their legal right to buy physical items and create digital backups?
>They make meth-heads look scrupulous.
My local meth head checks in on my family every 2-3 months, because when she had fled from hospital post surgery, and added some meth to some morphine, we gave her new clothes and a safe place while we convinced her that the ambulance service wasn't run by Satan. She's good people. Anyway if she wanted a whole bunch of digital books I would help her torrent them like a responsible person.
I find this argument weird because you or I are below the threshold of being targeted over book torrents these days.
MS ain't some altruistic company having core mission the good of humanity, as they proven across decades.
I have first hand knowledge of household name companies with billion euro IP who use OpenAI and Athropic products quite liberally. They have WAY more lawyers than I do and their sole job is to keep the IP safe. They wouldn't sign a deal with even a slightest whiff of the IP being used to train anything.
And if it happens, the penalty for breach of contract would have so many zeroes it'd be enough to buy a country.
It can be very hard to prove in court that your data was used for training
It very hard for a company to do anything without leaving some kind of paper trail behind that can be discovered in court. Not without crippling their own operations by simply refusing to digitise or write down anything.
Lots of companies filled with people working very hard to obscure their shady practices have been hoisted by their own internal docs. Just look at any major Uber, Google, Apple, Microsoft lawsuit. Do really think Anthropic and OpenAI are gonna be better at destroying their paper trail before the lawsuit starts?
corp usually have all sensitive stuff ttl deleted for this reasons.
Because, almost certainly, it's not your source code.
Did they? No, it worked out fine. Same with hosting your application at AWS instead of on owned hardware in a locked cage at the local co-lo data center.
Although of course the same argument was made against co-locating in a data center! Long ago an engineer carefully explained to me that no serious business would put their data in a co-located data center since the data center operator could just plug in a hard drive and take it all.
It turns out that contracts do actually mean things, and businesses want to do business with each long term. If you’re at a tier where Anthropic contractually commits to not train on your data, I would just sign and move on.
Still does not address ecological concerns, of course.
I don't think I can tolerate its writing style anymore. Reading Claude output is starting to cause actual psychological harm. I have tried many ways to get it to stop writing in its stupid punchy linked-in marketing-team voice, and I can't.
Is there a model out there that sounds sound this awful? It's like rubbing sand into the folds of my brain.
Often I'll tell it to summarize what it said only because it's providing too much detail, not that it's using esoteric language or weird claudisms.
No idea why people are still using Claude models.
My impression is they started on Claude Code and never tried Codex or other harness+model combos
As a side note, I think we are about too dismissive of computers and the Internet as “not real life”, that stuff written in text boxes by humans are just sticks and stones, etc. And now that we shouldn’t anthro-po-morphize or whatever the LLMs. But I do think, at least for some of us, that there is no way to avoid certain impacts that even non-personal text can have on us. It can feel alienating to interface with five different people in order to figure out how to solve a problem or navigate some burocracy (think Kafka). Well just interacting with one single LLM can induce that same feeling. The long-winded replies and the uncanny ways to miss the context.
Disclaimer that for my own needs I would be happy if the whole AI thing crashed and burned post-haste.
I’ve tried so many things and I have to beat Opus 5/Fable 5 over their head every time. With Opus 5, the failure to actually improve the writing is borderline comical. Both models have been building up tomes of memories on top of my core rules, all to be ignored.
Nothing sticks mid-writing! The only lever is to ask to revise after the fact, a particularly futile proposition for anything non-trivial with Opus 5.
In contrast, GPT 5.6 Sol Max absolutely obeys my edicts to write well. I selected the concise writing style and its default writing is really not bad, but tighten it up with a standing AGENTS.md order and it obeys.
The downside is that even at Max, it’s not as good as Fable or Opus 5 xhigh at writing code. Larger work and the 256k context’s forced compactions cause it to lose track/fidelity of critical details. Review passes are essential.
Another reason I’m considering leaving Anthropic are the sporadic refusals… Fable printed a Markdown body with a hex dump to inspect for trailing white space and line endings, boom, denied due to `reasoning_extraction`. Had to tell Fable to never do hex dumps. Asked it a few times “what do you think is typically done for this?”, got another `reasoning_extraction` error. It forces you to lose a whole turn of work when you have to press Esc/Esc to “retry” the last prompt… except if that prompt was mid-turn, you’re losing your entire turn.
The recent BashFirst experiment is hella nuts, where it prefers writing bash over Edit/Write tools in auto mode. WTF Anthropic?
I pity those stuck with Claude in an enterprise/work setting.
still don’t think anthropic models are worth the money
Sol gave me a straightforward bulleted list, easy to scan and read. Opus 5 though....it gave me a solid 7 paragraphs of how it worked. I read through it and yeah it nailed the same points Sol did but the output was way harder to read.
One of them writes better by default. One of them doesn't. That's the thing.
With enough steering, I can get a cheap low-capability model to do things correctly in most cases as well. But why bother?
You can tell it to write high quality code, to test things, and to come up with a proper rollout plan of a feature. Or it could just do it by default.
I know where I'm putting my money in that case.
I really love CC and have been using it exclusively for a year now but reading Claude's verbose and semantically-obfuscated writing style is wearing me out and I'm planning to move to another provider.
I hope they fix this. It's terrible.
Fable is much better; still much too verbose, but at least I don't have the feeling that I am being charged for gratuitously added verbiage.
I have been changing my processes and the roles of my agents to avoid interacting with Claude Code as much as I can.
> Calculation done. Potential confusion: sq km vs acres. Provide both. Output.
Other models would have written paragraphs dancing around the idea. Glimmer barely does sentences.
It doesn't really sound human - which is a plus in my book. I want my robots to talk to me like they are robots.
So stingy. Even with OpenAI's recent usage troubles, they're still so much better than Anthropic it's not even funny. Re-ran the code review with Sol as a benchmark and it turns out Sol's performance is within 70%-90% of Fable's. Anthropic's still got the best model, but what does it matter if I can barely use it?
I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter.
There is nothing as good as Fable, not even close.
I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues).
It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job of managing the number of agents to keep them within the 5 hour budgets (I think I had to restart it twice).
After 18 hours I ran a /simplify, /code-review, /simplify cycle which went for another 6 hours.
2 billion tokens (mix of Opus and Fable), 24 hours of continuous coding and a bug free outcome. It would have cost $2000 at API prices and worth every cent.
Fable's ability to keep other models on track while working on these long horizon goals is so much better than anything else.
Far from neutered, I've never had a cyber refusal, and Fable's English is actually readable (unlike Opus 5).
As an aside: while I hate reading Opus 5 English it still is a noticeably better model than Sol in my experience.
But I could handle losing Opus5 is I got Sol instead. But there is nothing even close to Fable.
in your instance, anthropic may decide, arbitrarily, to stop 'autonomous rebuilds / refactors and ports' because they could pose some alignment/rights/etc risk to whatever slop their philosophers dream up while they're out eating $200 avocado toasts. then you can't do the thing anymore.
fable is good, absolutely. agree it roasts Sol which is, comparatively, a little receipt-hunting jack**
but now imagine being an enterprise, and having another organization not only taking your workflows and baking it into your models, but then deciding they can arbitrarily cut you off.
when you can instead own your data, use an agnostic provider, and get better results (through model combinations), it will take 1-2 quarters to figure it out.
the main reason anthropic is killing it is because they really do understand the enterprise development experience and lifecycle and have built products and have a sales-team that can deliver.
business-model and vibes-wise they have lost all goodwill in the past 6 months, and that momentum will be quite hard to regain.
Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.
This is very true, especially vs Opus 5.
> tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.
The strength here was less the code quality and rather the long horizon task tracking.
This was a very large task - I was chatting with the maintainers and we estimated 4 to 6 months work over multiple phases for human coding.
Fable is able to handle that long goal, with incremental steps along the way, handle the verification and course correct when it finds a problem.
I think the larger context helps here some, but the strength of the model on this specific thing is notably better.
I'm not alone in noticing this. https://www.primeintellect.ai/research/nanogpt-speedrun shows Fable is able to manage a run nearly 1/3 longer than Sol (8.7 days vs 6.1 days). In my use cases Sol is much closer to Opus 5 though.
I think the rewrites are the main story for llms in code (hot take). Writing greenfield code at the seams also something which might work well.
This is true, but only for certain tasks. Even as a Fable fanboi, Sol is much better at Fable for some non-programming tasks: Fable for life-planning tasks is miserable because it keeps adjudicating rules, where I've found Sol to be insightful and warm (characteristics I'd previously associated with Anthropic models).
We're probably less than a year away from all the frontier models being so good at everything for day-to-day use that it doesn't really matter which you use, which is going to seriously fuck up the business models of all of these companies except the infra companies.
Sol in my eyes is powerful, but it over engineers so much, that its actually a liability. Where as Opus 5 is slightly under develops but you then can give it a small push for what is missing.
I rather have it under develop and i as the human in the loop, can correct/enhance it. Vs the models that adds so much, to the point that your going "dude, stop!". Remember, removing code for a LLM is way, WAY more difficult then adding it.
My main issue is with Sol is that its designed to over engineer without thinking why its doing something. Great that you security harden 1000s of lines of code, but ... nobody will ever get to that code. Its that lack of intelligence is where the model becomes a issue for me.
Its easier to have less code and then do security audits, with you approving what needs to be changed/hardend.
It may simply depend on the developers their mindset. Some folks just want the models to do everything for them, and performance or code bloat means nothing to them (forgetting that this bloat over time makes future LLM work more expensive).
Its funny how everybody has their own opinion for what model is better, when in reality its more about that model fits your own development style better.
That's 1.3M tokens per minute. I suppose you mean input tokens? This would not have cost $2000 at API prices because you would have cache hits.
Zero Data Retention, for the uninitiated readers in this thread.
The no-ZDR is clearly to permit surveillance. I would be shocked if NSA wasn't all up in these SOTA model providers' systems.
Never subscribe!
It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
Translation between languages.
That value dwarfs all programming value that can be had. Economically, culturally, scientifically, spiritually.
I compared insurance quotes last week. Needed cover for 2 brands, 1 company. Opus jumps up and down saying both brands need listing on the policy schedule. Human broker said not.
I told opus and it's the usual "thanks you're right" bollocks because it bothered to read in more detail and found that all business activities are covered.
Using them for coding makes it easy to self check its work (assuming those pieces of work are "verifiable").
As proofreading will still need to happen, what do you think the appetite for lawyers is to do this kind of work? Do you think it will drive fees down significantly? Empower younger lawyers at firms who probably are the ones doing this checking for the partners? (Or will that just create a further divide).
I'm genuinely asking as I am not in law but all my family is and it's nice to see someone here that's thought about the impacts in that space.
Will it drive down fees significantly? Doubt it, there's not enough pressure on them, and the industry is resistant to change. Firms don't want to make a big deal about using it because clients will then ask why they're not getting a discount.
Will it empower younger lawyers? Not as much as I'd like. I'm very fortunate to have an employer that lets me use Claude Code for my work (to a limited extent). I think for 99.9% of lawyers it's not an option available to them, through a mix of concerns around AI usage and concerned IT departments. There could be great benefits, but it would rely on having to break out of traditional private practice which would make it difficult to get enough work.
I think in the near term LLMs will have much the same impact on the legal industry as it has on the software industry.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
If you're actually willing to pick up old DC inference hardware on the cheap you can already do a lot at home. The new hardware not so much admittedly but if the Chinese models keep getting better and running on less hardware... I don't see a reason to be quite so pessimistic.
HN's take that they "kill their own company" seems wildly out of touch.
Truth is, Sam and Dario and Elon are all terrible and so are their orgs. Leaving any one of those companies for any of the others is wild.
Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)
Me: why did you add HTMX?
Opus 5: you asked it twice.
Me: quote the exact sentence(s) where I asked it.
Opus 5: I can’t because you didn’t.
So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
Check this comment of mine for a better explanation of this: https://news.ycombinator.com/item?id=49413353
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...
But that would require people to think! And the marketing is they don't need to do that...
Last consumer (ish!) product that required people to learn something to use it was PalmOS with Grafitti, wasn't it?
This work was fraught with bugs, a large portion of which came down to disagreement of what was coded v/s what was written. Even if you had airlines sign off sentence by sentence exactly what you wrote down in English, that's too open to interpretation.
People don't appreciate how the same sentence can be read five different ways by different people (or the same person on different days). We had to structure our documents to be closer to pseudo-code than to English to get any meaningful consensus on the definition.
Once the government stepped in, their 5th generation was effectively killed. They either had to lock it up (Mythos), neuter it (Fable), or leave people with the perception they were overpaying for a weaker model (Opus 5).
Hopefully Anthropic has learned to anticipate this risk and has a plan for rollout of their next model that plans for capricious ad-hoc regulation.
Among other things, "we have stolen the collected works of your culture, now help us grow so that you can all lose your livelihoods and become our serfs" has to be one of the absolute worst marketing approaches in history.
(Or perhaps vulns that NSA/etc discovered and has been keeping it private; as they're known to do).
I feel there's a lot that's unaccounted for, and the whole "AWS team reports a 'jailbreak' that is just 'review this codebase'" story doesn't add up.
I wonder if there were some parallel construction going on, and if at the same time, the NSA started losing the exploits they had because it was getting patched.
But this isn't what happened? It was only a couple of months ago! How can we be getting this completely backwards already!
Fable was aready what they were selling to the general public, and that is what the US government stopped them selling (not Mythos!)
Anthropic tuned the already existing classifier to handle the use case the government highlighted and then they went back to selling it.
They actually loosened the restrictions on ML-programming using it too.
I use Fable as much as I can and I've never had a cyber refusal for it.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
I'm just doing this as a hobby for fun and adding features to a program I use. I am trying to make the most of the boost they gave me for August. My account is max pro. (5x more than pro). I am also trying to make the most of this thing in the event that these business models do not pan out.
Computing technologies are relentlessly deflationary. If the value of their TAM wasn't a fraction of a manual process they replace, they wouldn't have a productivity advantage. And the amount of TAM per unit often declines over the life of that product category.
I would be unsurprised to find investment in data centers to be 10X what was really needed. The same goes for where the LLM S curve starts to flatten.
Then all the mythos warnings happened, then the fable access issues started.
At that point, I just wanted to try something else because I didn't want to feel vendor locked by anthropic.
Then I tried codex on the $100 plan. I've never been happier. It does everything claude can do, without any token limit issues or anything for me.
I don't see myself going back to claude anytime soon. I also highly recommend folks to not vendor lock yourself to a model. You need to try multiple models for yourself and see what's possible. There should be zero loyalty to any frontier lab because of how fast things change and how volatile model access, uptime and customer experience can be with these things.
This is just a case of competition working: Fable (and Opus 5.0) just aren't as good as Anthropic believes, and there's no switching cost, so everyone's moving away.
So far not looking back to Anthropic. They lost me.
outside of coding scenarios, anthropic is just not quite the best choice for regular folks.
depending on the cloud/office stack used by the company, the userbase is better poised to use the native ai offering than moving to anthropic and connecting everything with them. same goes for third party guys trying to tie it all together.
not to discount them or their features, but beyond small companies or startups, we don't see the same levels of adoptions now nor maybe in the near future.
those asking us to build genai solutions for them also do not choose anthropic first over the main model offered by their current CSP. anthropic made itself available across the board finally, but it was already too late for most projects planned out for the upcoming quarters. also, we don't use the absolute biggest model nor instantly migrate to the latest one. the insane price jumps do not help with the same. especially when we quote something that may no longer be true down the line.
aside: good luck to the FDE folks hired by the frontier labs. i wonder if they would really get to handle large company-wide projects in practice.