Excessive token usage in Claude Code
github.com
github.com
The agent now becomes unresponsive at times and needs a reload, which really breaks flow. More frustratingly, the context limit seems to fill up much faster on the same project that was working fine just days ago. Nothing major changed on my side, so this feels like a backend or token allocation shift.
Highly encourage people having issues to do /context and start removing heavy things. It's usually some sprawling MCPs they rarely use, or huge CLAUDE.md files they generated or cargo-culted from someone else.
I'm not suggesting these are the only ways to hit the limits, it's just (so far) almost always the answer when someone hits the limits doing something that I wouldn't expect to be problematic.
The final nail was them offering a $50 credit toward overage use that within a half hour of enabling maxed out and began digging into. It's become almost predatory now, and I have no way to quantify the actual usage I'm getting from it other than it burns now at an alarming rate.
Since I've stopped using Claude, I ultimately landed on Codex where for my usage, where I'm easily getting 4x less quota usage from it than Claude for the same period of heavy use. I keep it as a backup now if Codex gets stuck on something, but I'm annoyed enough to stop paying all together.
Run /context.
Observe that your MCPs are killing a sizable chunk of the usable context window.
The utility here is that it'll break down exactly which MCPs are consuming how many tokens, just in tool descriptions. Then you can decide if that's worth it to you, even if you continue using Codex or OpenCode or etc.
Yeah.
What's the 'news'?
I’ll believe it when I see actual facts, e.g. actual token counts (which is relatively easy to capture if you use mitmproxy or something like that).
For all I know this guy has a 5000 line CLAUDE.md
I should update my notes.
That’s the implied question. It’s an issue on a piece of software. There’s more than 10,000 of them.
Why is it news?
It's hard for Anthropic to cater to both sets of users with one model.
Both hobbyists and professionals are understandably frustrated that tokens are being consumed quickly without justification, or at least in ways that seem entirely avoidable.