Claude Code May–August 2026 weekly limits promotion
support.claude.com
support.claude.com
My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements. And as we see from some of the Chinese models, they can quickly distill frontier models with the tax of being slower and more token-guzzling, while retaining most of the quality. The real differentiators are becoming speed and efficiency, which translate to cost and user velocity more than incremental capability improvements.
It's great that frontier models can solve complex math equations, but bread and butter LLM usage (where the money is made) has already shifted from "I need the best always" to "what solves my day-to-day problems quickly and consistently". Fable usage as a percentage is flat-lining. We are already at the point where output quality is negligible. What wins going forward is cost, speed, consistency and the compounding effects of "softer" improvements to the harness.
gpt 5.5/5.6 goes further on its own much more often than opus 4.8/5 does. codex capped ~300k context when claude does 1m.
I don't feel codex is saving tokens, and result is usually not as good imo.
That's configurable in codex.. but there is a higher cost/usage to using it.
“No, not like THAT!”
When I checked, all that credit was gone, I still wasn't into the next 5 hours, and all the agents had failed, returning nothing.
I didn't even get anything for burning all that credit. If I had paid for it, I'd be very, very pissed.
One fine day I was like, oh my weekly quota resets soon, I will kick off an expensive bug hunt. It launched parallel things and burned $10 in like a moment.
It’s scary. I turned off extra usage after that.
"Ultracode. When a system-reminder confirms ultracode is on, that opt-in is standing: author and run a workflow for every substantive task by default. The goal is the most exhaustive, correct answer you can produce — token cost is not a constraint."
My boss uses opus and gets good results when I use it always burns tokens. The other way is also true too I get great results with sol/terra but my boss does not.
This has been a recurring problem for me - as a security engineer catasrophization is a fundamental skill to finding vulnerabilities in complex systems, but makes me to conservative when building. For some of my leaders they are much better at saying 'good enough, ship it, and fix it later'.
By contrast, Claude Code's bias to make assumptions of reasonableness about underlying systems has proven to be immensely frustrating over the last month or two, both personally and at work. I've wasted days on "that was my mistake. I've been reporting numbers on the old architecture because I hadn't enabled the new one in the config" both at work and home. It's immensely frustrating.
But here we are. Wrestling with energetic idiots in model form, wrangled by over-specific harnesses that struggle to stay off of deranged side-quests.
What a time to be alive!
Isn’t the race between Chinese open-weight models and the others more decisive for the future?
once we have a bit more memory fab capacity and the insane bottomless investment in AI giants realizes there is a bottom, local hardware will catch up with model performance to the extent that centralized inference will be downgraded to special cases or for orgs that find it cheaper than buying expensive hardware
but really most power users are going to have 1 TB unified RAM and local models that will do well enough
for light office use you can still have a cheap laptop and a claude subscription
Granted it did make me think about my biases and to lean into more future facing inevitabilities. But you've nailed it, for Boris and Anthropic, they are betting that coding is a solved problem in the sense that any person can one-shot any random idea and the output will be in some abstract sense "good". And then at what cost and toward what end?
As someone who does near 100% of my coding via LLM these days, i still find that for anything complex i am still looking at and thinking in terms of code. Im still quality checking and steering at some interval via code. And im still not sure how or whether i can replicate that level of thought without still dealing in code at times.
“In one notable technique, their prompts asked Claude to imagine and articulate the internal reasoning behind a completed response and write it out step by step—effectively generating chain-of-thought training data at scale. We also observed tasks in which Claude was used to generate censorship-safe alternatives to politically sensitive queries like questions about dissidents, party leaders, or authoritarianism, likely in order to train DeepSeek’s own models to steer conversations away from censored topics. By examining request metadata, we were able to trace these accounts to specific researchers at the lab.”
— “You are an expert data analyst combining statistical rigor with deep domain knowledge. Your goal is to deliver data-driven insights — not summaries or visualizations — grounded in real data and supported by complete and transparent reasoning.” (variations appearing 10s of 1000s of times)
https://www.anthropic.com/news/detecting-and-preventing-dist...Does Dario have the same relationship with the truth as Sam? (Their companies pirate books and develop products that compete at some level with those books’ authors, so obviously neither are that wonderfully trustworthy, so maybe “no evidence” meant you don’t believe this evidence rather than you weren’t aware of it. I would understand and respect your lack of belief!)
You break you pay .
That’s why in the western world you got book publishers . They have lawyers etc .If you prefer Chinese way : laugh at anyone who is taking about copyright , whether that a book , a model , or manufacturing secrets in eu , us or whatever sure .
A family business of my friend in eu was destroyed due to stealing until every single bolt position design of their system by China . They had no legal protection .
Keep calling rule of law vs no such rules an American exceptionalism .
It’s economically lucrative for everyone in that chain to attack again and again ai vendors , unless they are in China . In this case they can only cry .
This is evidence only if you're an American exceptionalist.
“These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts…”
and think “oh, there’s alleged evidence of distillation“.https://en.wikipedia.org/wiki/American_exceptionalism
Seymour Martin Lipset, a prominent American political scientist and sociologist, argues that the United States is exceptional in that it started from a revolutionary event. He therefore traces the origins of American exceptionalism to the American Revolution, from which the U.S. emerged as "the first new nation" with a distinct ideology, and having a unique mission to transform the world. This ideology, which Lipset calls "Americanism" but is often also referred to as "American exceptionalism", is based on liberty, individualism, republicanism, democracy, meritocracy, and laissez-faire economics; these principles are sometimes collectively referred to as "American exceptionalism".
As a term in political science, American exceptionalism refers to the United States' status as a global outlier both in good and bad ways. Critics of the concept say that the idea of American exceptionalism suggests that the U.S. is better than other countries, has a superior culture, or has a unique mission to transform the planet and its inhabitants.
Like, when I consider the Tuskegee Syphilis Experiments or MK Ultra, or Flock, I have bad things to say about them but American exceptionalism specifically wouldn’t cross my mind. Corporations anywhere playing a card close to the chest or even lying… superiority, exceptionalism, aren’t my top reference points.everyone else uses it as "Americans love to criticize everyone else, while pretending their farts dont stink"
Perhaps partially covered in the “Critics of the concept say…”
Certainly weird to think the US isn’t chock full of problems (even if historically some stuff has been neat, like bringing together cultures and developing nifty tech in Silicon Valley).
I've jumped between copilot, claude, gemini and chatgpt since the start of the year. chatgpt wasn't even worth looking at early this year.
Anthropic has the smarter models for sure, and seems to be default in corporate. However, the amount of budget you get with GPT as a user is much better, the harness feels more polished, and the models are faster. They are also much nicer to work with, I can just read the output for the most part. With claude I get pages of text and need to skim to find where the actual information i need to care about lies. So much more cognitive overhead.
Sol is smart enough for anything I've thrown at it, it's not one-shotting like fable, but I'm more willing to actually go back and forth with it, and it's likely producing better output to keep a human in the loop rather than trying to solve the world independently and making multiple incorrect assumptions.
I think GPT sees the market changing and is correctly repositioning themselves. Anthropic is down the wrong road, and if they don't correct course quickly I'm sure many of those enterprise contracts will start pivoting.
I'm not sure how they'll survive their creditors tbh
The collapse of the massive overvaluation and circular investments will be the undoing of several major tech investors and companies that have a serious contribution to the things in tech we don't like. It will be bad for the economy, but it will also weaken the abhorrent control that those companies have over regulation in the United States.
It's gonna be a rough time, but it's pretty clearly necessary.
The only big problem is that if that collapse happens under the current administration in the US, they will be utterly incompetent to respond to a real domestic crisis, since they have been virtually unable to do anything without creating more problems.
I was just discussing this with a co-worker yesterday. I would really like a model (or harness?) that worked with me instead of for me. Walk me through its choices and decisions, let me correct it and guide it along. I would be way more confident in it's output, I would be more familiar with the changes that are being made, and it would make reviewing the final code way easier since I was making the decisions along side it. I'm sure it would also reduce the "brainrot" we're all going to experience the more we hand work to these models.
I find it unlikely that there's some fundamental property of OpenAI's models' "personality" or style which Anthropic (or any other serious AI firm) wouldn't be able to match if they wanted to.
As much as I wish you were right, everything about software economics for the last 30+ years has favored _less efficient software_. Traditional hardware has been optimized for traditional software for decades and we still see bloated software win consistently. LLM hardware is at the start of its cycle, with abundant low-hanging fruit to conquer--I would expect the pro-bloat dynamics to weigh even more heavily in the LLM space than in the traditional software space.
They're both pretty damn competent and more important, Extremely Fast! I find the speed more useful than trying to be a hundred percent complete on every task. The big intelligent models screw up all the time as well, but I have to wait twenty minutes to three hours to find out.
GPT 5.6 is especially tenacious and seems to want to solve every bug in edge case 1000% all the time. Sometimes that's what you need, but a lot of times you're just trying to move fast and figure out what the product is.
I'm just happy there is capital behind the efficiency path because I am confident we can get positive value on that path.
Open weights, obviously beneficial. Local compute when not necessary, seems like it'd be significantly worse for the environment?
I think because more devs cant run things locally, we are not getting the "open source gains" that you would get by having the network effect of more eyes on the work. Theres still little nuggets of gold in there though (eg antirez's DS4 project). But once more people can get their hands on hardware, I think the flood gates will open. And once that happens, I dont necessarily think people will run locally. But that will enable fungible service providers which we already see, but at a much larger scale.
To a lesser degree, open weights also leads to an eventuality of these two companies not being able to sustain their capital investments. Not to mention their pricing models are unsustainable as they currently stand. So they are being eroded from the outside-in, while also trying to fight the internal pressure to continue to raise prices/margins.
EDIT: I actually think the harness is more of a future for these big players. And theres real value in providing a high quality harness. So maybe they move to SAAS for claude code in the future. Who knows
I do believe it though. While the subsidies of the frontier providers are nice, they can't continue forever.
Most ordinary people don't care that much about privacy and if there is a solution that just works well enough to cover their needs and costs just fair enough to be affordable, they are going to use that. I feel like this is why fast food is such a thing. There are far better food options out there that either require a little bit more effort or money. Your post on Linkedin is akin to someone claiming in 1960's America that home-cooked burgers are the future, despite McDonalds gaining ground.
It is also why crypto wallets and the like never took off, people don't want to give up the convenience of a bank account, understandably.
It's hard to put a number on it, but even accounting for all the time in meetings, talking with stakeholders and developers etc I'm over 10% more productive overall. I earn substantially more than $2000/month, the ROI is there
It's only expensive compared to the currently very strong offering from OpenAI. Or other models - my hobby projects are all on DeepSeek
Perhaps that’s true, but it’s a price tag that’s a lot harder to swallow for many orgs than $200/mo, and would require some hard justification for how your increased productivity contributes to the business bottom line.
I’ll agree that you can probably do that with hard numbers. I am skeptical that most $200/mo users could.
At the end of a long session it will start saying stuff like “There are smoke tests on the foundation-gates that are left for the cutting seam checks on these domains, which is genuinely your decision”
Never before the past couple months have I ever not been able to understand wtf it’s even saying lol
I was working on a project recently that required the attribution of a data source, and it added to the "licenses" page of the project something like "We use <blah> and per their terms we owe an acknowledgement of attribution to you, the user."
So in addition to the stuff it says in a session there's gobbledy gook that it prints out in copy as well.
The unnecessarily jargon filled language was already pretty bad, but the latest Opus/Fable also just didn't seem to go far enough when asked to look into something, often leading me in circles.
It's the better experience. The limits are way higher (I almost never burn through my $200/m plan), and the output is better than Opus 4.8 (Opus 5 is completely unusable for me).
Not doing crazy multi agent swarms, but have yet to hit any limits during pretty intense weekend sessions.
That's not bad (and nor is Luna!) but Opus/Sol are just massively better especially once you start doing agentic tasks.
See for example https://x.com/zainhas/status/2085567832268656755
It's clearly not right if you compare it to API prices. It's even less right if you compare to a SWE salary.
Yes not using the full $200 means you aren't tokenmaxxing, but that seems about the most you can conclude from it.
"From May 13, 2026 through August 19, 2026, your weekly usage limit in Claude Code is 50% higher."
Once on a sandbox, realizes nothing works in its sandbox, then again outside the sandbox.
Fucking hell.
It still starts off in the sandbox anyway. :/
What about errors compounding and any on-the-go decisions being made being the wrong ones? That kills anything long form, no matter how good and detailed your initial plan is - there will always be something along the way.
The models would need a way to identify a difficult problem and apply max reasoning there themselves and cruise through everything else at a lower level.
As soon as it gets annoying enough to switch to another provider your Claude code tokens drop right off.
Double plus good, eh!
"Your subscription will auto renew on Aug 20, 2026."
But I doubt they will remove it, just like they didn't remove Fable from the subscription. There's simply too much competition.
After they perma-extended Fable into including it into the max plan, I'm willing to give them the benefit of the doubt with this one.
"We hope to make this a permanent change to our plans, but strong demand for our models means that capacity may be tight over the coming weeks. We’ll keep you posted as things develop."
(nobody uses claude anymore, it's too over-subscribed)
Sideline LLMs have free compute and offer cheap prices to draw in crowd. Becomes flavor of the month LLM.
Back to step one.
It should be pretty clear by now that token prices are predominately a function of available compute.
https://www.bloomberg.com/news/articles/2026-08-17/anthropic...
I think the better question is how much more than $65B/year revenue they need to cover what they are spending on capex and model development. I would bet money their revenue in the next year is over $75B (vs $65B), but also that their amortized costs exceed their revenue.
I suspect it's more likely they will spend the extra money on hardware and infrastructure (data centers) either directly or via suppliers.
I wish I could use my Anthropic sub with it but I heard you get banned, but at least you can use it with any other subscription or model.
For anything that makes money, it feels like a huge step down in productivity, even with the downsides of agent-assisted code.
The max account is for home projects where I'm basically making some common tools but tailored to myself (diet tracker, note app etc).
It's been great but across work and home it's just too much.
I wish inline completions would have gotten better but it seems no effort is being put into it and it is frozen in time right now. Thankfully, the stuff i use this code for is simple and small scale. I can’t even begin to imagine how actual programmers are feeling about this.
Edit to add: As a 3 month experiment, it has been positive. I have gotten things to work that none of the people I have hired and paid over the years took care of. And I myself have been too busy to prioritize. It's a useful tool, but only a fool would think it replaces the judgment of a professional. Not yet anyway.
Like, stop toying around with token limits and just focus on more efficient models.
Once the local models are good enough we are so abandoning these elephants.
I'm gonna walk as soon as possible.
They raise their prices too high, a lot of customers will still buy but be unhappy about it, it's burning goodwill for money. If they drop their prices too low they have overwhelming demand.
They could be making money hand over fist for all we know. We don't really know how much compute is being used, nor how much it cost them.
This happened to me, so far two months without any progress with Claude support trying to resolve it. Chatbot support got no response, email support got a single response after a month saying it had been passed on to another team to resolve. Mostly just silence.
Why should I pay you money , and have slower model and less intelligence? Oh yeah and I do not care about limits with cursor ultra at all . Unlike with you .
But CC took some getting used to, I also prefer the Cursor interface (the legacy "chat" pane).
- I actually rarely hit the token limits using my $20/mo web plan (switch chats when the current one feels "bloated" due to huge context and therefore the token drain), which even covers some of my coding sessions using Claude Code. - The visuals and "interactive" elements are still not introduced in GPT's web UI, unlike Claude's. - The nail in the coffin is the fact that Opus 4.8/5 is just BETTER at doing my RTL verification review jobs than GPT 5.5.
I am not glazing the model; it's just a personal preference for getting the job done...
Edit: If you're mad I'm light on the details I have provided some in the replies.
I'll ask it to do something and it'll say, I tried, but I couldn't do it over and over again or some variation of.
But it doesn't do that at the start of my subscription, so...
Because just calling something bad does not add a lot to the conversation. It's not thoughtful, interesting, or good.
The Claude desktop app is also widely panned, as I mentioned, and for me this mainly is due to general UX and a poor remote control interface. Codex's connected machine support is top notch.
I also mentioned the value of the Codex resets!
Everyone has different experiences with these things (for example, I've never experienced what you describe), and "$X is bad" is not conducive to thoughtful discussion.
One thing I do like about Claude is that the normal (non-Code) chat interface supports MCP, whereas ChatGPT basically does not.
Opus 5 is fine for me and works better and faster on low and medium than higher effort on prior versions. Same as 5.6 Sol compared to 5.5 or 5.4.
I have a Claude Code and OpenAI subscription so that I can use Opus/Fable/gpt-5.6 as I please, and the models are often catching things the other models missed. So much that I would significantly weaken my workflow if I dropped one subscription.
My best workflow at the moment is to create the initial plan with Fable (before review/revise-cycling with other models). From my own testing it seems slightly better at arriving at high-level ideal solutions after sweeping the whole project, projecting future needs, then coming up with good trade-offs like "by construction" correctness.
While mostly subjective, maybe the closest objectivity I have here is noticing fewer revision cycles needed with Fable-initialized plans.
I mean i think if a developer has a good handle of the code the difference is marginal .
Unless we 100% offload the thinking to the model and act like a prompt manager. Maybe
Even then, it's kind of a wash these days between the sota models, and we're talking about maybe a 10% performance difference or something. But every once in a while there's the experience of one model spinning its wheels on a bug/repro/issue while another model comes in and one-shots the solution.
I am asking because in my personal projects after a while they becomes a giant messy ball of wires and i basically trust the model to untangle it for me , by the time it untangles properly, I run into my token limits.
You can swap out "architectural simplification" with performance opportunities, bugs, correctness, etc. I get the orchestrator agent to then itemize it all into a file where I can keep track of which ones I've implemented.
The results are pretty astounding. I run these right before my weekly limits reset for each subscription and the findings will dictate the secondary tasks I get done during the week.
It's definitely token-heavy. I'm on the $200/mo Claude Code sub and the $100/mo Codex sub.
But it's pretty clear to me that software engineering is more or less solved and all you need is enough patience + tokens to get what you want. I think 20 years of engineering experience more lets me save on tokens rather than unlock things nobody else can build.
An example of the scope of one of my AI-engineered projects is a iterm2/ghostty-like terminal app that implements its own pty session, parsing, rendering. It's almost 2000 commits right now.
That said, I have a specific workflow that isn't just a blind "ok now make it so a screen can be split into panes". I have a plan phase focused on coming up with ideal invariants and such. But I'm not sure anymore how much of that is useful vs just yoloing a solution and then paying technical debt in sweeps, like garbage collection.
But the example you give is of a terminal for which there are copius examples in open source code. How hard is it really for a pattern matching machine to do that?
If I was doing it I would start by forking an existing repo and I might even say then that "software engineering is more or less solved since the open source revolution"
But I don't work on things like that.
Even if it weren't, if you're capable of explaining the context and constraints of your problem, then a modern LLM with effort=high will generally come up with a solution that's worth starting with because it's well-reasoned.
Moreover, you can start with the solution and then course-correct based on future information because refactoring is trivial with an LLM, yet human projects often ratchet into a local optimum because refactoring is too expensive.
I don't think "it's been built before" does as much work as it seems. I didn't fork a project. The LLMs reasoned about how to build the project from scratch using trade-offs that made sense for my needs, and they made reasoned, unsolicited deviations from kitty, xterm, and co, not just blindly doing what some ref impl did. Btw, it was still a lot of work because my project isn't just "kitty but swift".
Then the models went on to drive a well-reasoned incremental implementation of a system that lets me use the terminal running on my Macbook from my iPhone over tailscale with a decent scheme it came up with itself.
The point is that I don't really have to know how things work to build good software with modern models. LLMs can do things like read Linux source code and adversarially refine ideas such that the final idea is a good one. And my biggest influences on the project can be automated through the use of reusable markdown files.
Now, I'm at risk of downplaying all my years in software here, but I see the writing on the wall. It was only one year ago that I only trusted AI to do autocomplete.
While just simply trying many times independently gets you the improvement that is due to pass@k vs 1, you can get huge improvements if on top of that, depending on your setting, you find a way to ensure some stochasticity by perturbing tool calls, etc and running multiple instances.
The general theme is, embrace the stochasticity rather than the leaky abstraction on top of it.
With modern LLMs, investing in this kind of harness tooling is much more fruitful than hoping for the best from the model.
While many basic instances of this are built in to the popular harnesses (much of cursors higher-than-usual success rate with older models was due to really excellent context mgmt), you can never beat one that is optimised for your particular codebase, infra and general setup.
Until last year or so, the context management needed varied too much at too coarse a level across different models and even model instances, but now they are all extremely robust in a much higher % of contexts and are thus way more amenable to developing context management tools for, without needing to do a research teams worth of evals.
Custom evals and harnesses are thus extremely high ROI now. We are finding companies needing to do less and less tweaks and getting much fewer regressions (you should have reg tests in ur evals) with every new usecase and every new model.
It can be really simple to start with: change your grep/rg that it uses to a script that does in effect "rg $@ | shuf".
More complex examples are: giving different subagents different tools, randomly failing tool calls, truncating file reads randomly, having a small model invent N possible failure modes causing a bug and appending that to N prompts and starting subagents from each - this all forces each to pursue different paths. $example_specific_to_your_company_setup is highest ROI though, since most companies actual failure modes are dominated by idiosyncratic API shapes and retrieval quirks that no usual harness will bother modelling.
Also important IMO to not assign any meaning or semantically interpret the CoT as an acceptance mechanism (it is ok to use it as a rejection mechanism e.g if you see it plotting a sandbox escape whether it eventually emits the exploit or not is not something you want to hedge). We have to resist the temptation and ensure we only interpret tool calls, codegen, etc in our evals and only think of the cot as "some output that pushes the conditional distribution" which may or may not semantically match the typical preceding tokens of the desired tool call.
The whole thing is getting ridiculous.
If Wall Street sees a single outage or a tiny drop in usage, they won't be happy and will pressure Anthropic to take away the free tokens.
Better to reduce the limits now rather than to wait until Wall St. tells them to just to avoid a stock punishment.
I'm kind of surprised that you haven't been watching it with /usage to see where you're at. I've had a couple times that it ran out and lost what it was doing and had to start over, so I've gotten more careful.
And the time it burned $100 credit (that they gave us) and the fizzled with multiple agents, leaving me nothing of that work... Well, that was certainly instructive.
TIL about the /usage command... Thanks!
And I suspect you're pushing Claude a lot harded than I am.
> the whole OpenClaw thing imploded
has it really imploded... ?seems like just a month or so ago it was the new hotness [0]
[0] https://www.cnet.com/tech/services-and-software/from-clawdbo...
The future is models baked into directly into hardware, onto bare metal, serving 10k+ tokens/sec.
We can have our little models, they'll be serving a different customer.
And smarter people than me can probably find some extra-linear relationship between usage and required limits for uptime etc.
I replied, "never stop stopping!" and we had a stalemate, lol
Yay, as expected.
"Thinking" aka 'trust us bro!' without proof of thinking.
Making the model waste more tokens.
Advertisement: aka you pay to be advertised at.
Silently downgrading you and still faking models with the more costly tokens.
Intentional strategies to eat more tokens with no real gains.
Giving out almost-but-not-quite solutions that require another pull of the slo(t/p) machine.
Don't take my word for it, use a simple prompt of "create a shader showing [complex scene]" and see what the other models vomit out, including opus 4.6/4.8/5 and 5.6 sol (compare fable low effort with sol xhigh).
So yeah, different people get different quality. I wouldn't be surprised if it calculates a wealth level, willingness to upgrade, influence, usage levels, etc. and decides what models to serve.
It's extremely obvious to me they've nerfed all models on my 5 Max from Fable to even 4.6, but i've tried through enterprise with huge differences in quality, and i've seen other people experiencing this both on reddit and twitter.
Over the long term I think OpenAI will produce the better experience when it comes to model quality, harness quality, and availability. I have been using codex the past few months and never looked back.